Skip to content

"Trends and Information on AI, Big Data, Data Science, New Data Management Technologies, and Innovation."

This is the Industry Watch blog. To see the complete ODBMS.org
website with useful articles, downloads and industry information, please click here.

Oct 1 26

Influence Is Not Ownership: Anna Widenius and Kaj Arnö on Governance, Open Source, and the Future of MariaDB

by Roberto V. Zicari

“The purpose of governance is to make clear where authority comes from, how decisions are made, and how other people can acquire responsibility and challenge decisions… Influence is not the same thing as personal ownership of the project.”

— Anna Widenius, CEO, MariaDB Foundation

Q1. Anna, you recently published what you called a “love letter” to the people who made it possible to document the governance processes that MariaDB Server has relied on in practice for years — making explicit what was already working but largely invisible. Why does making governance visible and inspectable matter so much for an open source project, and what specifically was at risk of being lost if these processes had remained undocumented and understood only by a small group of long-standing contributors?

Anna Widenius: For me, the central issue is that open source should not require insider knowledge.

MariaDB Server has had a mature and disciplined development process for a long time. There are experienced maintainers, established areas of technical responsibility, rigorous review practices, and well-understood ways of resolving technical questions. MariaDB plc in particular has invested enormously in building and sustaining that engineering capability.

What was less visible was the project-level map of that system.

If you were already deeply involved in MariaDB development, you understood who carried which responsibilities and how somebody gradually earned greater technical authority. For somebody approaching the project from outside, that pathway was much harder to see.

That is where transparency becomes important. A healthy open source project should make it possible for a capable contributor to understand what responsibilities exist, what standards are expected, how decisions are made, and what it takes to progress from contributor to committer, reviewer and eventually maintainer.

So the risk was not that MariaDB lacked professional engineering processes. Quite the opposite. The risk was that an effective system could remain more legible to the organisations and individuals already deeply involved in it than to the next generation of contributors we want to attract.

Documenting governance makes that system inspectable and transferable. It builds trust, makes succession easier, and gives external contributors a clearer basis for deciding whether they want to invest seriously in the project.

I sometimes think that open source code without inspectable governance is only half open. The source tells you what the software does. Governance tells you how you can become someone trusted to help shape what it does next.

So I see this as a sign of maturity: taking a development culture and governance practice that has evolved over many years and making it explicit enough that people outside today’s core group can understand it, participate in it and eventually help carry it forward.


Q2. The new governance framework makes one principle explicit: authority comes from the responsibility a person carries within the project, not from their employer. That is an important and deliberately stated principle — but it is also a genuinely difficult one to uphold when, as you acknowledge, most of the current maintainer responsibilities are carried by people working for MariaDB plc. How do you ensure that principle is not merely aspirational — that it actually shapes decisions when the interests of the Foundation, the commercial company, and independent contributors diverge?

Anna Widenius: I think the first thing is not to pretend that employer relationships are irrelevant. They are not.

MariaDB plc employs a large number of people who spend their working lives developing MariaDB Server. It would be very strange if those people did not have significant influence over the project. They have accumulated enormous technical knowledge, and many of them carry substantial ongoing responsibility.

The important distinction is between influence that comes from doing the work and authority that comes automatically from being employed by a particular company. An engineer’s technical authority within the project comes from what they contribute, what they maintain, what they review and what responsibilities they are prepared to carry over time.

To me, the real test of employer neutrality is not whether MariaDB plc engineers have a lot of influence. Given the amount of work they do, of course they should. The test is whether someone outside MariaDB plc who makes comparable contributions and takes comparable responsibility can earn comparable authority.

That is precisely why documenting the pathway from contributor to committer, reviewer and maintainer matters.

Different participants will naturally sometimes have different priorities or technical views. MariaDB plc may have a strong preference based on its customers and product strategy. The Foundation may be seeing a different need across the wider ecosystem. An independent contributor may propose something neither organisation originally considered. Technical questions still have to be resolved through the technical governance of the project, based on the merits of the proposal and the responsibilities of the people involved.

A strong commercial company investing heavily in upstream development is an enormous asset to an open source project. It means having engineers whose full-time work is to develop the product, maintain complex areas of the codebase, review other people’s contributions, test changes, resolve regressions, respond to production requirements and carry technical responsibility over many years. That kind of sustained professional engineering capacity is extremely difficult for any open source project to build and retain.

MariaDB plc provides a great deal of that capacity to MariaDB Server, and the project is substantially stronger because of it.

The Foundation’s role is to make sure that this strength exists within an open project structure: that participation remains open, that the path into greater technical responsibility is visible, and that contributors outside MariaDB plc who are prepared to make comparable contributions and carry comparable responsibility have a genuine path to earning comparable authority.

That, to me, is the point of employer neutrality. It does not require diminishing the influence of the people doing most of the work. It means making clear that influence ultimately follows contribution, expertise and responsibility, and that the same principles apply to everyone.


Q3. Kaj, you have observed that most open source governance frameworks are discussed only when something goes wrong — and that MariaDB’s step was to make what already worked visible and open to inspection rather than inventing a new structure. From your years of experience at MySQL AB, Sun, and now MariaDB Foundation, what are the most common ways that open source database governance fails silently — not in a visible crisis, but in a slow erosion of contributor trust, decision quality, or technical direction — and what early warning signs should a foundation board be watching for?

Kaj Arnö: Silent failure of open source database governance may happen in several ways. Through neglect, active and passive. Through ignorance, lack of attention. Through missing expectation setting, too little communication. Through allocation of only junior resources to community contributions. Through too slow response times by senior resources. All of these come as grey zones and all of them are areas where we have learned a trick or two over the years, moving into ever lighter shades of grey.

A foundation board who makes such mistakes is not likely to be listening for early warning signs. They are most likely to not even understand what they are missing out on, and sit there with only a handful of contributors and a backlog of pull requests that are slowly decaying, sometimes withering, other times festering and becoming increasingly hard to merge. So for a Foundation board, PRs languishing in the queue are probably the best early warning to look out for.

And to be honest, that’s exactly what happened in our case. The remedy in our case was the agreement with MariaDB plc, in yet another example of plc strategically committing to the health of the open source project, to ring fence time for reviewing community contributions in all developer sprints. As a result, we’re now in much better shape than before. 


Q4. The relationship between the MariaDB Foundation and MariaDB plc is one that the open source community watches carefully — because it represents a model that many projects are trying to navigate: a foundation that stewards a technology in the public interest, alongside a commercial entity that depends on that same technology for its business. Can you describe concretely how the roadmap development process works between the two organizations — who proposes what, how priorities are set, where the Foundation has genuine influence over technical direction, and where it does not?

Anna Widenius: I think it is important to separate several things that are sometimes bundled together under the word “roadmap.”

MariaDB plc has its commercial and product priorities and decides where it wants to invest its engineering resources. Those priorities are informed by customers, by production experience, by product strategy and by the very substantial engineering organisation it has working on MariaDB Server.

The Foundation has a different but complementary field of view. We hear from community users, hosting companies, cloud providers, distributions, application ecosystems, universities, tool vendors, independent developers and other organisations building around MariaDB. We also look at the long-term health of the Server, ecosystem gaps, contributor experience and areas where important work may not yet have an obvious commercial owner.

Other contributors and companies may have priorities of their own. All of those streams can feed into the same upstream project.

So I do not think of the MariaDB Server roadmap as a single document produced by one organisation and handed down to everybody else. It emerges from several sources of requirements, investment and technical initiative. MariaDB plc is by far the most important of those sources because of the scale and continuity of its engineering investment, but it is not the only source.

The Foundation’s influence comes primarily through participation and enablement. We can influence the roadmap where the community demands for a feature, identify ecosystem gaps, bring together people who should be working on the same problem, improve infrastructure and testing, sponsor or incubate work that strengthens the wider ecosystem, and help external contributors navigate their way into the project.

One area I am particularly interested in is extensibility. MariaDB already has a remarkably rich architecture for plugins, storage engines and other ways of extending the Server, and some of the most interesting opportunities now come from people and companies outside the traditional core development group.

That is exactly where open governance and technical architecture meet. If we want other people to build on MariaDB, they need both technical surfaces that allow them to innovate and confidence that there is a transparent path for engaging with the project.

To me, that is one of the strengths of the model: commercial product investment, community priorities and external innovation can all feed the same upstream project. These requirements don’t need to be conflicting and, above all, do not need competing roadmaps.


Q5. Customer and user feedback is one of the most important inputs to any software project’s roadmap — but in open source, the channels through which that feedback reaches the development team are often opaque. How does MariaDB Foundation ensure that the needs of organizations running MariaDB in production — particularly smaller users who are not Gold Sponsors and do not have direct access to MariaDB plc’s engineering team — actually influence the project’s direction rather than being systematically filtered out by the priorities of those with the loudest voices and the largest contracts?

Anna Widenius: This is one of the reasons an independent Foundation is useful.

MariaDB plc has direct access to the requirements of its customers. Those are extremely valuable signals, particularly because many of those customers operate MariaDB in demanding production environments and at significant scale. Customer requirements often reveal problems, performance needs and operational realities that are difficult to discover in any other way.

The Foundation adds another layer of visibility.

We are increasingly connected to hosting companies, cloud providers, e-commerce platforms, developer-tool vendors, observability companies, universities, independent developers, technology partners and users who may never become customers of MariaDB plc. We also try to create deliberate channels for listening beyond the organisations we already know: our annual MariaDB user survey, conversations at open source and database events, community meetups, contributor discussions and the many individual conversations that happen around those activities.

That gives us a different kind of signal. You could say that MariaDB plc often sees extraordinary depth, while the Foundation can sometimes see breadth across an ecosystem.

One user asking for something is useful information. Twenty hosting companies independently running into the same operational problem is a pattern. If developers, platform providers and application vendors are all telling us that the same piece of interoperability is missing, that becomes a very strong signal. The annual survey gives us another way of testing whether something we are hearing repeatedly in individual conversations is actually visible across a much broader population.

A major part of the Foundation’s role is therefore aggregation. We can see needs that might otherwise remain scattered across hundreds or thousands of individual users and bring those patterns into conversations with maintainers, contributors and MariaDB plc.

Aggregation also means going beyond counting requests. To make a signal actionable, we need to understand the workload, scale, operational requirements, economic impact and whether organisations are prepared to test or adopt a proposed solution. That allows us to bring maintainers, service providers and cloud platforms evidence they can use for real investment decisions.

Sponsorship gives organisations a more structured relationship with the Foundation, and I think that is entirely legitimate. Sponsors support our work, and naturally we invest time in understanding what matters to them. Their requirements become another source of information about how MariaDB is being used in the real world. Technical decisions, however, still go through the project’s normal governance and review processes.

The wider principle is that a healthy project needs multiple feedback channels. Commercial customers are one. Community users are another. Contributors are another. Surveys, open source events and the ecosystem around MariaDB give us still more.

The real value comes from putting those signals together. A large enterprise customer can reveal a very deep production requirement. A group of hosting providers can reveal a horizontal ecosystem problem. An independent developer can identify an innovation nobody else was looking for. A survey can tell us whether something that looks important from a handful of conversations is actually widespread. The broader and more transparent those channels become, the better the project can understand the world it is actually serving.


Q6. Kaj, you mentioned that the MariaDB Foundation Board meeting covered topics ranging from the Oracle MySQL Contributor Summit and ecosystem interoperability to AI skills, sponsorship strategy, and AI-assisted security analysis. That is a striking range for a single board meeting — and you publish fairly detailed minutes precisely because you believe open source deserves open governance. What is the most important and least-discussed topic from the perspective of the broader open source database community that is not yet getting the attention it deserves at the board level — and what would it take to put it on the agenda?

Kaj Arnö: We do try to cover “all” topics, at least to touch upon matters in order to set expectations for later deep dives, when we have more time. Lately, we have been indeed been able to cover a fairly wide range of topics. Our plugins, i.e. our extensibility strategy is a highly technical issue where ecosystem perceptions still sees us behind PostgreSQL, for no particularly good reason. This needs board level attention, and I’m sure it will bubble up fairly soon by itself, as other issues get resolved. Developer mindshare issues also deserve more attention, as do issues related to communicating our roadmap and functionality to the right channels. We keep hearing about users surprised at our feature set,”I thought you were just a fork of MySQL”, not understanding the results of years of work on top of MySQL. It breaks my heart to see us having a great Vector offering, when it comes to both versatility, performance and ease of use – yet we haven’t found our way to the right developer minds and forums in order for us to be a default choice. Those are key matters to be put on the board agenda.

Q7. The database market has never been more fragmented — or more dynamic. We have relational databases, document stores, graph databases, vector databases, time-series databases, NewSQL systems, lakehouse formats, and now AI-native data stores, all competing for developer attention and enterprise budget simultaneously. From the perspective of two people who have spent decades in the database industry: is this proliferation of database types and systems genuinely serving users and organizations well, or is it creating a fragmentation problem that makes it harder to build and maintain reliable data architectures — and where do you see the market consolidating, and around what?

Anna Widenius: I think we have probably overcorrected toward specialization in databases.

There are good reasons why specialized systems appeared. Different workloads have genuinely different requirements, and specialization has driven a tremendous amount of innovation. But every additional database also creates architectural cost: another operational system, another security model, another data pipeline, another consistency boundary, another skill set the organisation has to maintain.

So I expect some consolidation, but not necessarily consolidation into one giant database that tries to implement every possible workload itself.

The more interesting model is a strong, extensible core that can absorb new capabilities when they belong in the core, while also supporting best-of-breed engines and extensions for specialised workloads. That gives users a much broader range of capabilities without requiring a completely separate database platform, operational model and skill set for every new use case.

Vector search is a good example of the first path. A few years ago, it was easy to assume that vector workloads necessarily required an entirely separate category of database. MariaDB now provides native Vector capabilities as part of the Server itself. That is a case where an important new workload can become a natural capability of an established relational database.

Other workloads may be better served by specialised technologies connected to the same core. Analytics is an interesting example, where work connecting MariaDB with DuckDB shows the potential of combining MariaDB with a best-of-breed analytical engine instead of trying to recreate every specialised capability inside the Server.

This also matters when analytical requirements emerge inside an established transactional application. Organisations should be able to address many of those requirements without moving the operational workload to another database or creating an entirely separate data platform. MariaDB with DuckDB points toward a model in which transactional and analytical capabilities can coexist behind the same familiar database interface, while dedicated analytical systems remain available for workloads that genuinely require them.

This is particularly interesting for MariaDB because extensibility is not something we have to invent from scratch. It is deeply embedded in the architecture. MariaDB has had pluggable storage engines for decades, alongside many other plugin interfaces.

So the opportunity is not simply to keep adding more and more functionality to one monolithic database. It is to give users one strong and familiar core that can evolve itself where that makes sense, while also allowing specialised engines and extensions to serve very different use cases efficiently.

And that has an important open-source dimension. The most interesting innovation is not necessarily the innovation the core project itself predicted.

A healthy extensible platform combines a professionally maintained core with an ecosystem that can innovate around it. Maintaining a database Server over decades requires sustained engineering work: compatibility, reliability, testing, regressions, performance, security and all the less glamorous work that makes new capabilities usable in production. MariaDB plc contributes enormous capacity to that part of MariaDB.

Extensibility then allows the circle of innovation to become much wider. The core project does not need to predict every future workload or build every possible capability itself. It needs an architecture and interfaces that allow other people and companies to build useful things around it, while benefiting from the strength of the underlying platform.

I think that combination of a strong professionally maintained core and an open ecosystem capable of extending it is going to become increasingly important.

Q8. The MariaDB governance framework describes how contributors can grow into committers, reviewers, and maintainers — a pathway that exists in theory in many open source projects but in practice is navigated successfully by relatively few people. What are the real barriers that prevent technically capable contributors from moving along that pathway, and what has MariaDB Foundation learned from cases where the seven-day response clock you mention has failed to produce the momentum that was hoped for?

Anna Widenius: There are several barriers to becoming a contributor to a mature database project. The codebase is large, ownership can be difficult to understand, testing can be intimidating, and even technically excellent contributors can lose momentum simply because nobody responds quickly enough.

That is why we have been working on very practical things: making responsibility clearer, improving contributor documentation and testing, and setting expectations around response times. A seven-day clock can prevent silence. It cannot manufacture mentorship, but it can at least make sure that somebody who has made the effort to contribute is not left wondering whether anyone noticed.

But I think we also need to broaden what we mean by participation.

Historically, open-source projects have tended to think about the path from user → contributor → committer → maintainer. That path remains extremely important. We want more people eventually carrying real responsibility for MariaDB Server.

At the same time, a mature database needs people with very deep specialist knowledge who are willing to carry responsibility over many years. MariaDB plc employs many such engineers, and that continuity is enormously valuable. It is neither realistic nor necessary to imagine that every person or company contributing value to MariaDB should ultimately become a core Server maintainer.

Someone may build a storage engine, an extension, an integration or an entirely new capability on top of MariaDB. A company may want to build a product around the database, provide independent support and consulting, integrate MariaDB into a cloud platform, or operate it as a managed service while contributing improvements to the interfaces and capabilities it depends on. Those organisations are also part of the technical and commercial ecosystem around the project.

So one of the things I want us to become much better at is making MariaDB easy not only to contribute to, but to build on.

That requires good documentation, stable interfaces, testing, discoverability and clear ways for extension authors to interact with the core project. Governance matters here as well, because people invest much more confidently when they understand how decisions are made and can see a realistic pathway for deeper engagement with the project.

For me, that is a broader definition of openness. The source code is open. The path into responsibility is open. And increasingly the architecture itself should invite people to build things around MariaDB that none of us could have designed centrally.

That gives us a much larger ambition than simply increasing the number of people committing patches to the Server. It means creating an ecosystem in which professional core engineering and innovation around the core reinforce each other.

One measure of a genuinely open project is whether several independent companies can confidently build support, services, tooling and managed offerings around the same upstream technology. Transparent governance and stable technical interfaces make that investment possible and help the project reach users through many different commercial channels.

Q9. Looking at the challenges ahead for MariaDB as a project and as an ecosystem — AI workloads requiring new database capabilities, the rise of vector databases, increasing competition from cloud-native managed database services, and the ongoing challenge of sustaining open source contribution at scale — what is the challenge you are most concerned about that you believe the broader open source database community is not yet taking seriously enough?

Anna Widenius: MariaDB is a tool that at the same time is served by AI and itself serves AI. Integration of a tool into AI usually means making it easy for the harnesses to develop code for that tool. Making vibe coding efficient. In our case, we are also an infrastructure component of AI, based on our Vector offering. The RAG applications enabled by MariaDB make it possible for users to create AI applications that scale, and that save tokens by delegating work to “normal” MariaDB indexes, instead of consuming oodles of AI tokens. I say “normal” as we’re talking vector algorithms, such as HNSW (which stands for “hierarchical navigable small world”), where a smart design of the app avoids the costs related to large contexts.

The challenge is the complexity of the AI world and the speed at which tools change. New harnesses and new LLM models are hard enough for many of us to navigate, and the best usage of skill files and MSP servers is changing rapidly. Architecting a scalable app in such a situation is of course a challenge to be concerned about, particularly as new innovations may quickly obsolete best practices. The trick is to balance the ambition between quick wins and the sustainability of a long-term app.

Q10. Anna, you share the surname Widenius with Michael “Monty” Widenius, the creator of both MySQL and MariaDB — one of the most significant figures in open source database history. Kaj, you co-founded MariaDB plc with Monty and spent years as CEO of the Foundation he created. The question neither of you may want to answer publicly: how much of MariaDB’s identity, direction, and community culture is genuinely shaped by the governance framework, the Foundation board, and the community — and how much is still, honestly, shaped by the vision, preferences, and personality of one person? And is that a strength, a risk, or both?

Anna Widenius: Oh, no Roberto – this is a question I absolutely want to answer. In fact I am extremely grateful for the opportunity to do it!

Of course Monty has enormous influence.

Pretending otherwise would be both inaccurate and rather silly.

He created MySQL. He created MariaDB. He has spent decades thinking about relational databases, and a great deal of the architecture, technical philosophy and culture of both projects carries his fingerprints.

But there is an important distinction between enormous influence and MariaDB being driven by one person.

MariaDB today is developed by a substantial engineering organisation at MariaDB plc, together with maintainers and contributors across the wider project. Product direction, engineering priorities and technical decisions emerge from the work of many people: plc leadership and product teams, engineers responsible for different parts of the Server, maintainers, customers, the Foundation, and contributors bringing ideas and requirements from elsewhere in the ecosystem.

Monty is a particularly influential participant in that system, and he is also a member of MariaDB plc’s leadership team. He challenges technical assumptions, proposes ideas, argues about architecture and sets an extraordinarily high bar for what he believes MariaDB Server should be capable of. His influence sits within a much broader product and engineering organisation, alongside other executives, technical leaders and engineers with substantial responsibilities and expertise of their own.

His ideas are reviewed, challenged, implemented, tested and very often argued about by people with deep technical expertise and strong opinions of their own.

That distinction is important to me because governance is not a mechanism for averaging away exceptional people.

The purpose of governance is to make clear where authority comes from, how decisions are made, and how other people can acquire responsibility and challenge decisions. Someone with Monty’s knowledge, history and continuing contribution should have enormous influence. But influence is not the same thing as personal ownership of the project, and it does not mean that one person determines MariaDB’s product direction.

Opinionated, passionate discussion has always been part of the culture around Monty. He wants people who care enough about the technology to argue with him.

You see that in MariaDB today as well. The engineering organisation at MariaDB plc contains very strong technical personalities, as does the Maintainers Council and the wider contributor community. Governance gives that culture structure: responsibilities are explicit, decisions can be challenged, and authority can be earned by others.

So where is the risk?

The risk would be confusing the enormous influence of a particularly capable individual with a system that depends upon his consent.

Those are two very different things.

I very much want Monty to remain a major intellectual and technical force in MariaDB for as long as he wants to be. At the same time, the measure of a mature project is that it becomes stronger through many consequential people and institutions: strong engineering leadership, strong maintainers, new contributors, more organisations carrying technical responsibility, and clear ways of making and challenging decisions.

To me, that is the ideal outcome. We do not make MariaDB Server “less Monty” in order to make it more open. We make sure that the extraordinary contribution of its founder exists within a project and engineering organisation capable of producing many other people whose ideas can be equally consequential.

Kaj Arnö:  My answer is at the same time different from Anna’s, as it is identical to it. From the outside, it may seem that the quirks of the founders of a project direct most aspects of an Open Source project. To the extent it’s true, it shows on a cultural level. The importance of technical merit is there, and the sheer joy of winning an intellectual argument has always played a role in the MySQL and MariaDB universe. With MariaDB, arguing from a community ethics perspective became even more important than during MySQL AB times. In between during Sun Microsystems Inc. times, there was a detour into a preaching mode, where exporting “our” model to Sun was high on the mental agenda. But the really interesting thing is to see how Monty evolves. Old dogs don’t learn to sit, proverb has it. Yet Monty has learned quite a few tricks in the last, say, ten years: In communication, in expectation setting, in listening to other opinions. They say you don’t change much after you turn 35; Monty has changed his opinions on core engineering topics. To name just one, he has moved from begrudgingly accepting that others use AI, into vibe coding himself (not core Server code, but test cases and peripheral scripts). He was even happy to review and improve upon a vibe coded contribution to a fundamental new piece of functionality (JSON and BLOB support in in-memory tables). Monty remained Monty and found lots of things that absolutely needed optimisation, improvement and partly rewriting. But even if Monty’s vision remains intact, his personality and preferences adapt and evolve with the times we live in.

……………………………………….

Anna Widenius, CEO MariaDB Foundation

Anna Widenius is CEO of the MariaDB Foundation, where she champions open-source database technology and works to strengthen the MariaDB ecosystem. She brings extensive experience in open-source advocacy, marketing, operations, and strategic planning. Previously serving as Chief of Staff at the Foundation, Anna played a key role in driving strategic initiatives and community engagement. She is passionate about open source, innovation, and building vibrant communities around technology.

Kaj Arnö, Executive Chairman MariaDB Foundation

Kaj is a software industry generalist, having served as VP Professional Services, VP Engineering, CIO and VP Community Relations of MySQL AB prior to the acquisition by Sun Microsystems. At Sun, Kaj served as MySQL Ambassador to Sun and Sun VP of Database Community. Board member at Footbalance Systems Oy (Helsinki, Finland). Past founder, CEO and 14 year main entrepreneur of Polycon Ab (Finland).

Kaj is a co-founder of MariaDB Corporation Ab, and served on its Executive Team in several positions, most recently Chief Evangelist.

………………….

Follow us on X

Follow us on LinkedIn

Sep 30 26

Why AI Agents Struggle with Complex Software Failures: A conversation with Greg Law

by Roberto V. Zicari

“Code generation has become cheap. Understanding software behavior hasn’t.”

— Greg Law, Founder and CEO, Undo

Q1. Your new research suggests that AI has made writing code easier, but 79% of engineering leaders say release cycles are no faster than before. Why hasn’t the productivity gain from code generation translated into faster software delivery?

Greg Law: Because writing code isn’t the whole job. AI has become extraordinarily good at producing code, but for mission-critical enterprise systems such as DBMSs, somebody still has to understand what that code does, determine whether it’s correct, and generally take ownership and accountability for it and be responsible when something goes wrong. 

Modern coding agents are amazing but we’re a long way from being able to vibe code mission-critical infrastructure. AI can help with code review and understanding and debugging, but it’s not advanced at the same rate as it has for code generation, so the gap is widening and the bottleneck is becoming more and more extreme.

Our research found that engineers spend around 42% of their time debugging, compared with roughly a quarter actually producing code.

As an industry we’ve dramatically accelerated one part of the software development lifecycle without really addressing some of the much larger bottlenecks downstream. In some cases, we’ve made them worse. 

Code generation has become cheap. Understanding software behavior hasn’t.


Q2. The report argues that writing code was never the main bottleneck in software engineering. Has the industry been measuring AI productivity in the wrong place?

Greg Law: I think so. Lines of code generated is attractive because it’s easy to measure. But nobody gets paid for producing lines of code. We get paid for delivering software that works. As Bill Gate said decades ago, we shouldn’t talk about lines of code produced, we should talk about lines of code spent!

The more interesting measures are things like: How long does it take to get a change into production? How many defects escape? How long does it take to understand an unexpected behavior? What’s the mean time to root cause when something fails?

If I generate a thousand lines in seconds but spend two days working out why the system is behaving incorrectly, I haven’t gained very much. 


Q3. What changes when AI can generate code much faster than engineers can understand it? Are we creating a “comprehension gap” in software engineering?

Greg Law: Yes, and I think that’s one of the more important consequences of AI-assisted development that isn’t being discussed enough.

Historically, if you wrote the code yourself, you at least started with some mental model of what it was supposed to do. You might not understand every consequence of it, but there was a degree of inherent comprehension. There was intent.

With AI-generated code, that relationship changes. You can have perfectly plausible-looking code entering a system that nobody on the team really understands.

The survey found that engineering leaders estimate 35% of AI-generated code is not fully comprehended before it reaches production.

And the difficult failures aren’t normally caused by a line of code that is obviously nonsense. They’re caused by code that’s almost right.

Database systems are a good example. A change can be perfectly correct in most executions and fail only under a particular transaction ordering, thread interleaving or sequence of state changes. Those are exactly the failures that are hardest to reason about from source code alone, and hardest to diagnose when they strike in production.

Those are exactly the problems where looking at the source code isn’t enough. You need to understand what the software actually did.

So the risk is that AI increases the amount of software we’re producing much faster than it increases our ability to understand that software. 


Q4. The report says 80% of engineering leaders believe coding agents struggle to solve difficult problems in large-scale, complex codebases. Why? Is this fundamentally a limitation of today’s models, or a limitation of the information we give them?

Greg Law: Increasingly, I think it’s an information problem. The obvious reaction when an agent fails is: we need a better model. But that’s only half the story.

Imagine asking a very experienced engineer to diagnose a concurrency failure in a storage engine. You give them the source, some logs and the documentation, but you don’t let them run the program or inspect what actually happened.

They might come up with a very plausible hypothesis. But that’s what it is — a hypothesis.

AI is in exactly the same position.

Source code tells you what a program could do. It doesn’t necessarily tell you what it did during a particular execution.

My view is that continually throwing a more powerful model at the problem misses something fundamental. If the input is incomplete, you’re asking the model to guess. Give the agent better evidence and suddenly even a less expensive model can become dramatically more capable. 


Q5. You argue that agents need “runtime context.” Can you explain precisely what runtime context is, and how it differs from adding more logs, traces or other conventional diagnostic information?

Greg Law: Logs and traces are examples of runtime context. And agents are great at ingesting this runtime context and combining with the source code. But it’s exactly the same problem as human engineers have when looking at logs and traces – they’re incomplete. If you get lucky the information you need is in there. All too often, it’s not; the critical piece of info about what the system did in the lead-up to the error is just not there. 

For years Undo has provided human engineers with complete program recording of program execution, effectively 100% of the runtime context. We need to bring that to the agents too.

Not just a handful of log lines or a stack trace when it crashes. Capture the execution so that afterwards you can ask: Where did this value come from? Which thread wrote it? What sequence of events led us here? 

That’s the complete runtime context.

Logs are useful, but they’re selective. Somebody has to decide beforehand what to log. When the bug depends on something you didn’t log, you have to add instrumentation and reproduce it again. In many cases, if you knew to log it ahead of time you probably wouldn’t have made the mistake in the first place! (Where ‘you’ here can be a human or an AI.)

A recording changes that. Undo creates a deterministic, self-contained recording of execution that can be interrogated afterwards.

For an AI agent, that is enormously powerful because it can move beyond reading descriptions of the software and start reasoning about software behavior.

Static context helps an agent understand what the code says. Runtime context helps it understand what the code did.


Q6. How do you establish that an AI-generated root-cause analysis is correct? In mission-critical software, is producing an answer enough, or must the agent also be able to provide evidence that engineers can independently verify?

Greg Law: Producing an answer isn’t enough. Models are very good at producing explanations that sound plausible, but plausible and correct are two different things. 

If an AI tells me, “this looks like a race condition,” my next question is: Show me. Which threads were involved? What was the ordering of events? Where did the incorrect value come from? 

The interesting thing about recording execution is that the agent doesn’t just have more information with which to form a hypothesis. It also has evidence against which that hypothesis can be tested.

That changes the nature of the interaction. You can move from “I think this may be the cause” to “This value became incorrect at this point in execution; it was written here, by this thread, as a consequence of this earlier event.“

This is especially important with today’s AIs which suffer from being confidently wrong. It might give a confidently correct answer 4 times out of 5, but that fifth time where it goes down the rabbit hole due to being confidently wrong can be very expensive: a lot of time (and tokens!) are spent going down a blind alley, and eventually the agent just gives up and tells the human to fix it themself, often steering the human in the wrong direction first!

For mission-critical software like databases, networking systems, large-scale simulations, that distinction between assertion and evidence is going to become increasingly important. 


Q7. Your research report suggests that better context may allow less expensive models to solve problems that would otherwise require frontier models. Is the next major AI optimization problem therefore not simply “which model should I use?”, but “what information should I give the model?”

Greg Law: Very much so. There’s been an enormous amount of focus on model selection, but model intelligence is only one variable. The other variable is context.

A very capable model with poor context will perform worse than a less capable model with exactly the evidence it needs.

Complex debugging tasks can be extraordinarily token-intensive if the agent is repeatedly reading huge codebases, generating hypotheses and starting again, especially when it gets into the add logs, rebuild and rerun loop. Give it a recording of the relevant execution and you radically reduce the search space. 

Nearly two-thirds of the engineering leaders we surveyed said that the cost of higher-tier models was prohibitive for this kind of work.

So I think the next stage of enterprise AI isn’t simply going to be about buying access to a smarter model. It’s going to be about giving models better inputs.

In software engineering, runtime behavior is one of the most valuable inputs we’ve historically been missing.

In fact, our CTO recently presented data showing a cheaper model with runtime context outperforming a more expensive model working from logs. (ref. CppCon 2026 conference poster). This demonstrates clearly and quantitatively how rich runtime context makes AI smarter, faster and cheaper.


Q8. Looking three to five years ahead, do you expect engineers to spend less time debugging themselves and more time supervising autonomous investigations? If so, how does the role of the software engineer change?

Greg Law: It’s already happening. Many of our customers today have fully automated CI pipelines that go through triage, root-cause analysis, fix. Recordings are a key part of this to get the confidence that the diagnosis is correct and the fix works as expected. If we’re to realise the productivity promises of AI, it’s the only way: we need to apply AI to the whole SDLC.

When a test fails or there’s a production issue, an agent should be able to capture the failure, interrogate the execution, trace causality and present the engineer with the root cause and the evidence behind it, and then ideally suggest the fix. Some of our customers are already automating root-cause analysis in CI and, in some cases, production. 

That doesn’t remove the engineer. It changes where they spend their time: less on collecting evidence and reconstructing what happened, more on deciding the right architecture and broader engineering judgments. 

For me, that’s a far more interesting application of AI in software engineering than simply generating more code.


Q9. Your research breaks out findings by industry and suggests that data management teams are particularly concerned about the volume of AI-generated code entering complex codebases: 60% say engineers struggle to review it effectively and end up merging too many AI-generated errors, compared with 46% across industries overall. Why do you think this challenge is especially pronounced in database and data-management software? 

Greg Law: I think databases are a particularly unforgiving environment for this because correctness is so dependent on behavior over time. It isn’t enough for a function to look locally correct. You have concurrency, state transitions, recovery paths, memory behavior and interactions between subsystems, often over very long-running executions. A bit of state gets corrupted and everything carries on working just fine for some time until eventually the system falls over. And these things are deployed at such a scale that if something will go wrong one time in a billion then it’s going to happen all the time!

So AI-generated code can look entirely plausible in review and still introduce a failure that only appears under a particular runtime condition. That makes “just review the generated code more carefully” a pretty weak scaling strategy. 

Increasingly, the important question is not “Does this code look right?” but “What did this change actually cause the system to do?” And source code alone often can’t answer it, neither for a human nor for an AI.

Q10. The report finds that 82% of engineering leaders are concerned that the token costs of coding agents will rise sharply as Anthropic and OpenAI prioritize IPO economics. If runtime context lets cheaper models match frontier models on debugging tasks, does that weaken the case for always using the most powerful model and change how engineering teams think about model selection?  

Greg Law: I think it changes the economics quite fundamentally, because a bigger model isn’t always the answer. If you’re asking it to debug from incomplete evidence, you may just be paying more for a better guess. And it’s not just about cheaper models – if the AI gets the right answer first time, rather than repeated guesses and adding logging and rerunning, then it’s going to get to the conclusion using far fewer tokens.

Nearly two-thirds of the engineering leaders we surveyed said the cost of higher-tier models is already prohibitive for comprehension and debugging.

So the question becomes less “What’s the best model?” and more “What’s the least expensive model that can solve this reliably with the right context?” 

Resources

Research report: Overcoming the limitations of coding agents in complex software systems  

Agentic Debugging Live – Without a Safety Net, GenAI-X

…………………………………………………………

Greg Law is the founder and CEO of Undo. A systems engineer at heart, he has spent more than 25 years building and leading software teams working on complex, high-performance codebases where failures are costly and debugging is a bottleneck. His career spans Acorn, startups including NexWave and Solarflare, and ultimately the creation of Undo. 

Undo’s core technology grew from Greg’s firsthand frustration with traditional debugging. While at Acorn, he and co-founder Julian began developing a different approach: recording real program execution so difficult bugs could be diagnosed deterministically rather than guessed at. 

As CEO, Greg has taken Undo from an early technical breakthrough to an enterprise platform used by engineering teams that cannot afford broken releases, prolonged outages, or unresolved defects. He is also a regular conference speaker on software engineering, debugging, and the challenges of building reliable complex systems. 

Greg holds a PhD from City University, London, and lives in Cambridge with his family. These days, his coding mostly happens on long flights. In his spare time, Greg catches up on emails. 

Sep 23 26

Shaping the Intelligent Age: Klaus Schwab on Techno-Capitalism, AI Concentration, and Collective Intelligence

by Roberto V. Zicari

“The unit of responsibility must be at least as large as the unit of power.”

— Klaus Schwab, Founder of the World Economic Forum

Q1. In 1971, you published Modern Enterprise Management in Mechanical Engineering and argued that a company must serve not only shareholders but all stakeholders to achieve long-term growth and prosperity. That same year you founded the World Economic Forum to promote the stakeholder concept. Fifty-five years later, Thriving and Leading in the Intelligent Age extends that foundational argument into a world shaped by artificial intelligence. How has the arrival of the Intelligent Age changed — or deepened — your original conviction about stakeholder capitalism, and where do you see the greatest tension between the stakeholder model and the economic logic of how AI is currently being developed and deployed?

Klaus Schwab: The Intelligent Age has not changed my conviction; it has tested it more severely than anything since 1971. When I wrote about the enterprise’s obligations to clients, employees, investors, and society, I was arguing against a narrow model of ownership. Today the argument is no longer only about who owns the enterprise. It is about who controls the systems on which everybody else depends: the computing infrastructure, the data, the models, the platforms through which people find work, credit, information, and opportunity. I have come to call this configuration of power “techno-capitalism,” and it is organized differently from the industrial or financial capitalism that stakeholder theory was originally built to civilize.

The greatest tension is this: the economic logic of frontier AI rewards scale, speed, and concentration. A handful of companies now produce the overwhelming majority of the world’s most capable models, and the infrastructure beneath them (chips, data centers, energy) is similarly concentrated. That logic is not evil; scale can fund the safety research and the security that a fragmented industry could not afford.

But it is in tension with a stakeholder model that assumes negotiation among relatively balanced parties. A worker instructed by an algorithm, a small business dependent on a platform’s ranking rules, a creator whose work trains a model without consent. These stakeholders often have no realistic exit and no seat at the table. Stakeholder capitalism must therefore evolve from an appeal to enterprise conscience into an institutional architecture of rights, graduated duties, and remedy that is sized to match the scale of the power it is meant to govern. The unit of responsibility must be at least as large as the unit of power. That is the deepening the Intelligent Age has demanded of the idea I first set out fifty-five years ago.

Q2. You describe the Intelligent Age as one where intelligence — human, artificial, and collective — has become the world’s most powerful resource. That framing places collective intelligence alongside human and artificial intelligence as a distinct and co-equal category. What do you mean by collective intelligence in this context, and why do you believe it is as important as the other two — particularly at a moment when AI seems to be accelerating individual and organizational capability in ways that may actually be reducing the capacity for genuine collective reasoning and shared deliberation?

Klaus Schwab: By collective intelligence I do not mean the sum of many individual intelligences, human or artificial. I mean the capacity of a group, a company, a profession, a nation, an international community to reason and decide together in ways no single mind, however augmented, could achieve alone. It is what allowed the World Economic Forum, across five decades of crises, to convene competitors, regulators, unions, and civil society around problems that none of them could solve unilaterally. It is a genuine third form of intelligence, with its own conditions for success: trust, shared facts, structured deliberation, and institutions capable of holding disagreement without fracturing.

The danger is that we replace multiple personal interactions through AI. In other words, AI absorbs collective intelligence. This might make collective decision processes faster and more efficient but deprive them of qualitative reflections and subject them much more to failures in the algorithms. This is why I place collective intelligence beside the other two rather than beneath them: an economy of extraordinary individual and machine capability, without the institutions and habits of shared reasoning, produces fragmentation dressed up as progress. Cultivating collective intelligence in the Intelligent Age is not a nostalgic exercise. It is the only mechanism I know of that can direct enormous individual and artificial capability toward outcomes a society actually chooses, rather than outcomes that simply emerge.

Q3. You call for leaders to build what you describe as “strategic intelligence” — the ability to anticipate change before it disrupts. But the history of the Fourth Industrial Revolution, which you named and described in your 2016 book, suggests that even the most thoughtful observers consistently underestimate both the speed and the direction of technological disruption. What does strategic intelligence actually look like in practice for a leader who is genuinely trying to anticipate the next inflection point in AI — and where do you believe the conventional wisdom among business and political leaders today is most dangerously wrong about what is coming?

Klaus Schwab: You are right to press on this, and I will not pretend the record of forecasting, mine included, has been perfect. When I described the Fourth Industrial Revolution in 2016, I underestimated how quickly generative systems would move from research curiosity to general-purpose infrastructure, and I was not alone. Strategic intelligence, in practice, is less about predicting the specific technology and more about building an organization’s capacity to notice weak signals early, to hold several futures as plausible at once rather than committing to one, and to act on incomplete information without waiting for certainty that will never come. It looks like scenario discipline, dedicated resources for horizon-scanning that report directly to leadership rather than being buried in an innovation unit, and the humility to revise strategy in public when the evidence changes, rather than defending a forecast that has already been overtaken.

Where I believe conventional wisdom is most dangerously wrong today is in treating the coming disruption primarily as a question of which tasks AI will automate. That framing has leaders competing over efficiency gains while missing the structural questions: who controls the infrastructure and models on which entire economies increasingly depend, and what happens to accountability when responsibility for a harmful outcome is distributed across a model developer, a deployer, a data provider, and a platform, each of whom can plausibly claim compliance while no one accepts responsibility for the whole. Leaders who anticipate only the labor-market disruption will be well prepared for the wrong crisis. The leaders with genuine strategic intelligence right now are the ones asking who governs the systems, not only how fast the systems will change what their people do.

Q4. Your book calls for leaders to align innovation with ethics, empathy, and long-term value creation — but in practice, the competitive dynamics of AI development create enormous pressure to move faster than ethics can follow, and to capture value before governance frameworks can catch up. You have spent fifty years convening the world’s most powerful leaders at Davos, watching how those competitive pressures play out across industries and nations. What have you learned about what actually causes leaders to slow down and make the harder, values-based choices — and is there any realistic mechanism, whether regulatory, competitive, or cultural, that can reliably produce that outcome at the scale and speed the Intelligent Age demands?

Klaus Schwab: Fifty years of convening leaders through oil shocks, financial crises, and technological upheavals has taught me to be skeptical of appeals to conscience alone. I have watched extraordinarily thoughtful executives make decisions under competitive pressure that they would not defend at a dinner table. Exhortation rarely changes behavior at the pace or scale we need; incentives and accountability structures do. What actually causes leaders to slow down is some combination of three things: peers and investors who treat responsible practice as a condition of doing business rather than a virtue signal, regulation that is credible enough to be a genuine cost of non-compliance, and competitors who have made the harder choice public and durable enough that others cannot easily undercut them on it.

So no, I do not believe there is a purely cultural mechanism that will reliably work at the speed the Intelligent Age demands. What I believe can work is a graduated architecture: enforceable stakeholder rights for those most exposed to an AI system’s decisions, differentiated obligations that fall more heavily on the gatekeepers and providers of systemic infrastructure than on smaller challengers who cannot absorb the same compliance burden, and international interoperability around a limited set of principles: human oversight, transparency, effective remedy, so that responsible companies are not simply undercut by jurisdictions with none. This is beginning to take shape in instruments like the OECD’s due diligence guidance and the Council of Europe’s Framework Convention on AI. None of it is sufficient alone. Together, incentive, accountability, and interoperable standards can do what appeals to values consistently fail to do on their own — make the harder choice also the rational one.

Q5. You argue that in the Intelligent Age, human and artificial intelligence must work in synergy rather than in competition. That is a compelling vision. But it assumes a kind of partnership that current AI deployment trends do not obviously support — where the dominant commercial logic is to replace human judgment with AI judgment wherever the economics favor it. What does the synergy you describe actually require from organizations, from education systems, and from public policy that is not yet in place — and what is the most important thing that leaders can do right now to preserve the conditions for genuine human-AI collaboration rather than simply accepting substitution?

Klaus Schwab: You are describing accurately the default trajectory, and I do not want to understate it: left to pure cost logic, substitution will win more often than augmentation, because substitution is easier to measure and faster to realize on a balance sheet. Genuine synergy requires deliberate counterweights at three levels. Organizations need what I call an augmentation covenant, a negotiated, measurable compact attached to every major AI transformation that identifies which tasks will change, involves employees or their legitimate representatives in redesigning the work rather than simply announcing it, provides paid time to learn, and reports honestly on whether productivity gains are strengthening job quality alongside financial performance. Education systems need to treat continuous professional renewal as part of employment itself, with portable, recognized credentials, rather than assuming a young person’s initial training will carry them through a working life. And public policy needs mechanisms, collective bargaining in some countries, portable benefits and learning accounts in others, that ensure the intelligence dividend is broadly shared rather than accruing only to the owners of the machines. None of this exists at the scale required today.

The single most important thing a leader can do right now is sequence the covenant before the deployment, not after. Too often the transition compact for workers arrives once displacement has already occurred, as a remedy rather than a design principle. A leader who negotiates the augmentation covenant at the same moment the AI system is adopted – not after the layoffs – is making a structural choice for partnership rather than an after-the-fact gesture of goodwill. That sequencing, more than any statement of values, is what determines whether an organization ends up with genuine human-AI synergy or simply substitution with better public relations.

Q6. Thriving and Leading in the Intelligent Age is described as the first volume in the Intelligent Age Series — which suggests you see this as a sustained intellectual project rather than a single statement. What are the most important questions that this first volume cannot fully answer — the ones you believe will define the subsequent volumes — and how do you think about the risk that the technology will evolve faster than any series of books can track, making each volume obsolete before the next one is written?

Klaus Schwab: I conceived this as a series rather than a single book precisely because no single volume, however carefully argued, could do justice to a transformation of this scope. Thriving and Leading in the Intelligent Age sets out the orientation a leader needs: strategic intelligence, collective intelligence, the synergy between human and machine judgment. It deliberately leaves open the harder institutional questions that the later volumes take up directly: how a competitive enterprise structures itself and its governance when intelligence itself is the contested resource; how societies restore truth and trust when the information environment is increasingly mediated by systems whose incentives are not aligned with either; what longevity and an extended working life mean for retirement, care, and the social contract once both human lifespans and machine capability are changing simultaneously; and what universities, professors, and students owe one another when the premise of higher education, that expertise takes years to acquire and then holds its value, is itself under pressure.

On the risk of obsolescence, I hold two convictions that may seem in tension but are not. The specific technological detail in any book about AI will indeed be outdated within a year or two; I accept that, and I would rather publish something useful now than wait for a certainty that will never arrive. But the questions I am asking who controls the infrastructure, how responsibility is allocated when it is distributed across many actors, how gains are shared, what collective reasoning requires, are not questions about a particular model generation. They are questions about power, purpose, and accountability that recur across every technological transition I have lived through. A book anchored in those questions ages far better than one anchored in a particular capability benchmark. That is deliberately how this series is built.

Q7. You have engaged with every major global crisis of the past fifty years — oil shocks, financial crashes, pandemics, wars, and now the rise of AI — from your unique position at the center of global economic and political dialogue. Looking across all of those crises, what do you believe the Intelligent Age has in common with previous transformative moments — and what is genuinely unprecedented about it in ways that make the leadership frameworks developed for previous disruptions inadequate or misleading?

Klaus Schwab: Every crisis I have observed closely – the oil shocks of the 1970s, the financial crisis of 2008, the pandemic – shared a common structure. A shock exposed a gap between the speed of events and the speed of the institutions meant to manage them, and the resolution, when it came, required actors who normally compete or distrust one another to cooperate around shared facts and a shared sense of urgency. That is also true of the Intelligent Age, and it is why I still believe convening remains valuable: no single company, government, or discipline can govern this transition alone, any more than any one of them could have managed 2008 alone.

What is genuinely unprecedented, and where I think frameworks built for earlier disruptions can mislead us, is twofold. First, previous industrial transitions multiplied human physical power; this one multiplies, and in places substitutes for, human perception, analysis, prediction, and judgment itself, the very faculties earlier crisis-response frameworks assumed leaders and institutions would use to manage the crisis. Second, and this is the point I find most consequential, the governing power in this transition is disproportionately private and global from the outset, whereas oil shocks and financial crises, however global their effects, were ultimately resolved through instruments, central banks, national regulators, treaties among states that possessed clear public authority. Today, systems with quasi-public consequences are governed principally through corporate structures accountable to shareholders and, unevenly, to national law. Leadership frameworks built around the assumption that a sufficiently empowered state or central bank can eventually act decisively do not map cleanly onto a crisis whose infrastructure a state may not even be able to fully inspect. That is the genuinely new leadership problem of the Intelligent Age.

Q8. Your book is addressed to executives, entrepreneurs, professionals, and curious learners — a deliberately broad audience. But the leaders who most need to master change with purpose in the Intelligent Age may be those in the public sector — governments, regulators, international institutions — whose capacity for adaptive leadership has historically lagged far behind that of the private sector. What is your most honest assessment of whether public institutions are capable of leading in the Intelligent Age rather than simply reacting to it — and what would need to change, structurally and culturally, for that to become possible?

Klaus Schwab: My most honest assessment is that most public institutions today are not positioned to lead in the Intelligent Age, and that the gap is not primarily one of intention. I have sat across the table from ministers and regulators who understand the stakes as clearly as any technology executive. The gap is capacity: governments are frequently in the position of procuring, regulating, or depending on systems they lack the internal technical competence to fully inspect, replace, or govern on their own terms. A regulator who cannot evaluate a model except through what its developer chooses to disclose is not a regulator in any meaningful sense; it is a dependent customer with enforcement powers it cannot confidently use.

What would need to change is structural before it is cultural. Public institutions need enough internal technical competence — hired, retained, and paid competitively — to be what I call an intelligent buyer, regulator, and guardian of the public interest, rather than outsourcing that judgment to the parties they oversee. They need procurement and partnership frameworks with published mandates, disclosed decision rights, independent evaluation, and built-in provisions for review and knowledge transfer, so that public-private cooperation strengthens public authority instead of substituting for it. And internationally, they need to align on a limited set of interoperable principles, human oversight, transparency, effective remedy, because no single national regulator, however capable, can govern systems that operate globally by acting alone. The cultural shift, which does matter, is a willingness among political leaders to treat technical competence in the civil service as a strategic asset worth investing in during a period of comfortable growth, rather than a cost to be cut, because that competence is precisely what determines whether government reacts to the Intelligent Age or helps direct it.

Q9. The World Economic Forum has been both celebrated and criticized across its fifty-year history — celebrated as a convening force for global cooperation on shared challenges, criticized as a forum that brings together extraordinary power without adequate accountability, and more recently as a symbol of a globalist consensus that large parts of the world’s populations have actively rejected. As you look at the Intelligent Age and the leadership challenges you describe in this book, how do you think about the role that institutions like the World Economic Forum should play — and what has the rise of populism and democratic disillusionment taught you about the limits of elite-driven approaches to governing transformative technologies?

Klaus Schwab: I take the criticism seriously, and I would not be honest with you if I dismissed it as mere misunderstanding. Any institution that convenes concentrated power, even for constructive purposes, has to answer the question of who it is accountable to, and for many years that answer was not as clear or as public as it should have been. The rise of populism and democratic disillusionment has taught me something I should perhaps have internalized earlier: that convening is not the same as legitimacy, and that legitimacy cannot be borrowed from the credentials or good intentions of the people in the room. It has to be built through transparency about mandates and financing, structured representation for the workers, citizens, and communities affected by the decisions under discussion, and independent evaluation of outcomes, the same institutional safeguards I now argue must govern any public-private cooperation, applied to convening itself.

As for the role institutions like the Forum should play in the Intelligent Age, I believe it is narrower and more disciplined than it may once have appeared: a place where governments, companies, and civil society can build the shared facts and working relationships that crisis cooperation requires, but never a substitute for the democratic institutions, legislatures, courts, elected governments that alone hold final authority over public purpose. Populism’s core objection, at its most legitimate, is to elites making decisions with public consequences and no public accountability. The answer to that objection is not to convene less. It is to ensure that convening is explicitly subordinate to democratic authority rather than a parallel channel around it. That is a harder discipline than the Forum has always practiced, and it is one I think this era requires of every institution like it.

Q10. You were born in Ravensburg, Germany in 1938 — which means your first years were lived through the most catastrophic failure of human governance in modern history, followed by a lifetime spent building institutions designed to ensure that such failures could never happen again. You founded the World Economic Forum at thirty-three, wrote your first book on stakeholder capitalism the same year, and are now, at eighty-eight, writing a new series of books about how humanity should lead through the most powerful technological transformation it has ever faced. That is an extraordinary arc — from a child in wartime Germany to one of the most influential institutional architects of the twentieth and twenty-first centuries. What does the Intelligent Age look like from that vantage point — not the strategic analysis, but the personal experience of having lived through enough of history to know what is genuinely at stake when the tools of power change faster than the wisdom to use them wisely?

Klaus Schwab: I was a small child when the institutions meant to restrain power in my country failed completely, and I grew up in the physical wreckage of what that failure produced. I do not say this to draw a dramatic parallel between that catastrophe and artificial intelligence; they are not the same, and I am wary of anyone who reaches too quickly for the darkest historical analogy. But that early experience left me with a conviction I have never lost: that the tools available to human beings, whether industrial machinery or intelligent systems, do not by themselves determine whether they are used well or badly. What determines that is whether the institutions of restraint, accountability, and shared purpose keep pace with the tools. Germany in the years before my birth had enormous industrial and organizational capability and catastrophically weak institutions to direct it. That asymmetry, more than any single technology, is what I have spent my professional life trying to prevent from recurring.

So when I look at the Intelligent Age from this vantage point, what strikes me is not the technology itself, remarkable as it is. It is the pace. I founded the Forum at thirty-three because I believed, as a young academic, that the postwar economic order needed better mechanisms for cooperation than it had, and I assumed there would be decades to build them. The Intelligent Age does not obviously afford us decades. What I have learned across fifty years is that wisdom, the capacity to ask not only what we can build but what we should build and whom it must serve, has never moved as fast as our tools, in any era I have lived through. It does not need to move faster than the tools. It needs to move fast enough that the gap between the two never again becomes as wide as it was in the country of my birth. That is the only test that has ever mattered to me, and it is the one I am asking this book, and this series, to serve.

Qx. Anything else you wish to add?

Klaus Schwab: Only this. Readers will finish the Intelligent Age book series looking for a verdict on whether artificial intelligence will ultimately be good or bad for humanity. I do not think that is the right question, and I have resisted answering it throughout my career for every previous technological transition, because it has never had a single answer. The outcome has always depended on choices made afterward, in boardrooms, parliaments, classrooms, and communities, by people who had not yet decided how to use what they had built. The Intelligent Age will be no different. Its outcome is not predetermined by the technology. It will be determined by whether we build institutions, at the pace this moment demands, capable of directing extraordinary power toward purposes we actually choose. The ultimate test will not be whether our machines become more intelligent. It will be whether we do.

Resources

Thriving and Leading in the Intelligent Age: Mastering Change with Purpose

………………………………………………………………………………………….

Professor Klaus Schwab (1938, Ravensburg, Germany) is the Founder of the World Economic Forum. In 1971, he published Modern Enterprise Management in Mechanical Engineering. He argues in that book that a company must serve not only shareholders but all stakeholders to achieve long-term growth and prosperity. To promote the stakeholder concept, he founded the World Economic Forum the same year.
Professor Schwab holds doctorates in Economics (University of Fribourg) and in Engineering (Swiss Federal Institute of Technology) and obtained a master’s degree in Public Administration (MPA) from the Kennedy School of Government at Harvard University.
In 1972, in addition to his leadership role at the Forum, he became a professor at the University of Geneva. He has received numerous international and national honors, including 20 honorary doctorates.
His books include The Fourth Industrial Revolution (2016), a worldwide bestseller translated into 30 languages, Shaping the Future of the Fourth Industrial Revolution (2018), The Great Reset (2020), and Stakeholder Capitalism (2021).

……………………….

Follow us on X

Follow us on LinkedIn

Sep 16 26

Traceability Matters: Gopal Shankar on Opening Up MySQL Development for the Next Decade

by Roberto V. Zicari

“Would making this capability available help a meaningful part of the MySQL community build better applications or run better systems?”

Q1. MySQL has recently made available several features from the enterprise tier to the MySQL Community Edition. What is driving this direction?

The simple answer is that Community Edition needs to stay strong for the people who build, run, and depend on MySQL every day. That includes developers, DBAs, startups, large enterprises, and software vendors.

When we bring broadly useful capabilities into Community Edition, particularly in areas like observability, high availability, performance, and developer experience, the whole MySQL ecosystem benefits. Users get a better platform, and we get better feedback from real workloads at scale.

Recent examples include replication observability and Group Replication capabilities, OpenTelemetry support, the Hypergraph Optimizer, Profile-Guided Optimization, and enhanced JSON Duality View support in Community Edition.

I see this as a clear sign of Oracle’s long-term commitment to MySQL. The Community Edition is fundamental to MySQL. Enterprise Edition continues to address additional commercial support, security, and operational needs, but the success of Community Edition is essential to the success of MySQL overall.

Q2. When Oracle makes a feature available in Community Edition, how are those decisions made?

There is not a single checklist that applies to every feature. We start with a practical question: would making this capability available help a meaningful part of the MySQL community build better applications or run better systems?

From there, we look at technical maturity, operational impact, security, compatibility, documentation, and long-term maintainability. A feature has to work well not only in a carefully controlled environment, but also in the many different environments where MySQL is deployed.

The strongest opportunities are often in areas that improve the daily experience of using MySQL: understanding what is happening in the server, operating highly available systems, diagnosing performance problems, and reducing unnecessary complexity for developers. We will continue to evaluate those opportunities as MySQL evolves.

Q3. What does MySQL do better than competitors in community engagement, and where can it improve?

MySQL’s strength is the combination of a mature open source database, deep engineering investment, and an ecosystem that has been built over decades. MySQL runs important workloads for organizations of every size, so reliability, compatibility, upgrades, tooling, and operational simplicity matter deeply to our community.

We also see opportunities to improve. The best open source communities make it easy for people to understand how an idea moves from discussion to action. We want to improve that path in MySQL through clearer design discussions, better issue triage, more visible contribution paths, and more predictable review and feedback.

Code is important, but it is not the only meaningful contribution. Testing, documentation, benchmarking, production feedback, tools, and community education all make MySQL better. We want contributors to feel that these forms of participation are recognized and useful.

Q4. MySQL 9.7.0 LTS brings capabilities previously limited to MySQL Enterprise Edition—including JSON Duality Views, the Hypergraph Optimizer, and improvements to replication observability and HA behavior—into MySQL Community Edition. Which of these changes do you think will have the greatest real-world impact for DBAs and developers, and what should teams consider before adopting them in production?

I would separate immediate operational impact from longer-term developer impact. For DBAs running Group Replication, the replication observability and HA changes will probably have the fastest and broadest benefit. Better visibility into flow control, applier lag and throughput, unhealthy members, and primary election helps teams diagnose problems earlier and make failover behavior more predictable. This is practical day-to-day value, especially for teams operating clusters at scale.

For developers, JSON Duality Views may be the more consequential change over time. They let teams work with JSON documents while retaining relational integrity and a single source of truth. The Hypergraph Optimizer can also be significant for complex queries, but its benefit will vary more by workload.

Teams should approach an LTS upgrade with thorough validation. Test representative workloads, failure scenarios, replication behavior, upgrades, and monitoring integrations in staging first. For the Hypergraph Optimizer, compare plans and performance for important queries. For JSON Duality Views, validate the data model, update paths, permissions, and concurrency behavior. And for telemetry, make sure the collector, retention policy, and handling of potentially sensitive operational data are ready before turning it on in production. The Community Edition additions cover replication and HA behavior, telemetry, JSON Duality Views, and the Hypergraph Optimizer.

Q5. You were personally involved in designing JSON Duality Views. What problem does it solve?

JSON Duality Views solve a problem many application teams know well. Developers often prefer JSON because it maps naturally to APIs and application objects. But relational modeling gives them normalization, transactional consistency, referential integrity, and SQL.

Historically, teams often had to choose one model, or build and maintain their own mapping layer between application objects and relational tables. In some cases, they also ended up duplicating data across multiple systems.

JSON Duality Views let an application work with hierarchical JSON documents while the underlying data remains relational. The application can use the model that feels natural for the task, but MySQL still provides a single source of truth.

For a team, that can mean less mapping code and simpler synchronization. It does not remove the need for good schema or API design, and it will not fit every application, but it gives suitable workloads a much simpler way to combine document-style development with relational strengths.


Q6. How do you balance new capabilities with MySQL’s simplicity and reliability?

MySQL has earned trust because it is practical. People can deploy it, operate it, upgrade it, and troubleshoot it with confidence. New capabilities must preserve that experience.

We pay close attention to defaults, configuration, backward compatibility, documentation, and operational behavior. Early Access builds, LTS releases, compatibility testing, and upgrade guidance are the practical mechanisms that help us validate that balance before broad adoption. The MySQL 9.6 foreign-key work is a good example. We moved foreign-key checks and cascades into the SQL layer so that those changes are visible to binary logs and CDC tools, while preserving compatibility, validating performance, and providing a temporary `innodb_native_foreign_keys` fallback for staged adoption. A feature should be powerful when users need it, while preserving the straightforward core MySQL experience for everyone else. 

Not every user needs every new capability. Success is giving developers and DBAs useful new options while maintaining the stable, predictable MySQL experience that existing users rely on.

Q7. What does the more open community model look like in practice?

For a developer or DBA, participation should not begin only when they have a patch ready. They can discuss roadmap topics, share use cases, test Early Access releases, file actionable bugs, join GitHub discussions, contribute documentation or benchmarks, and participate in community events and contributor summits.

The important change is connecting these activities more clearly. If someone raises a good issue or proposal, they should be able to see where it goes next. Does it become a design discussion? A bug investigation? A request for testing? A roadmap input? That traceability matters.

Over the coming releases and community cycles, the community should see clearer guidance, more public technical discussion, improved GitHub workflows, and more structured ways to engage early. What matters most is whether people find the process easier to use and receive useful follow-up.

Q8. What changes are being made to improve the contributor experience, and how will you measure success?

The first improvement is clarity. Contributors need to know where to start, what information is needed, how review works, and what happens if a proposal is not accepted as submitted. 

We are working toward clearer templates, better-defined contribution paths, more visible technical discussions, and stronger links between issues, proposals, patches, and bug records. This should make it easier for contributors and users to follow the progress of an idea.

The second improvement is feedback. When a contribution needs further refinement, people should receive a clear outcome, and where possible, practical guidance about what to do next. 
We will look at evidence: response and resolution trends, time to initial triage, review cycle time, contributor growth, contribution quality, roadmap participation, Early Access adoption, and feedback from contributors. We also intend to share progress regularly, pairing timely acknowledgment with meaningful follow-through.The goal is not simply to collect more pull requests. It is to create a community process that produces better outcomes.

Q9. MySQL recently marked 30 years. What will define the next decade?

MySQL must remain the practical, dependable choice for the application workloads that matter most: transactional systems, cloud-native services, distributed applications, and data-intensive workloads.

That means continued investment in performance, high availability, observability, security, developer productivity, and operational simplicity. These may not always be the most visible areas of innovation, but they are the reasons people trust a database in production.

The other important aspect is community participation. MySQL cannot thrive for another decade based only on work from one company. The community can influence priorities earlier, contribute effectively, build tools and extensions, share operational knowledge, and see that its feedback leads to visible action.

By the time MySQL turns 40, I would like it to be known not only for scale and reliability, but also for a community that has a real and practical role in shaping its future.

Qx. Anything else you wish to add?

I would encourage people to engage with MySQL early and directly. Try the Early Access releases, share concrete production experience, bring specific use cases, and tell us where the friction is.

The most useful feedback is grounded in real workloads and comes with enough detail for us to act on it. MySQL has always evolved through the combined work of engineers, users, customers, partners, and contributors. We want the next chapter to be even more collaborative.

………………………………………………………..

Gopal Shankar, Director of MySQL Engineering, Oracle.

For over 20 years, I have worked at the heart of database engine architecture. Currently, as the Director of MySQL Engineering, I lead the organization responsible for the strategy, development, and roadmap of one of the world’s most popular database platforms. My expertise lies in the core internals of MySQL specifically kernel-level development, performance tuning, and scalability. I believe in solving complex technical challenges by prioritizing architectural simplicity and resilience. I have led the design of several important features, including the MySQL 8.0 Data Dictionary and Information Schema, as well as the recent JSON Duality feature. Additionally, I have helped architect the integration of foreign key handling directly into the SQL layer, effectively resolving long-standing trigger cascade limitations in MySQL recently. My goal is to deliver features that are high-performing, reliable and developer-friendly. I am focusing towards executing the MySQL roadmap and ensuring MySQL platform remains a powerful, solid foundation for modern applications. Beyond strategy, I stay connected to the kernel-level complexities tackling issues like high CPU usage, database corruption, and throughput bottlenecks. I am focused on fostering technical excellence and delivering a database engine that evolves with the needs of the industry.

https://www.linkedin.com/in/gopal-shankar-1b34664/

………………….

Follow us on X

Follow us on LinkedIn



Sep 9 26

Leaving the Bubble: Oliver Günther on Universities, Democracy, GenAI, and What Public Higher Education Owes Society

by Roberto V. Zicari

“As a university president, one asks oneself every day: What can universities contribute to solving today’s global problems? Their financial and staff resources may be limited, but they have the power of the word and the mind. Through their teaching, they ensure that future generations can deal with complex challenges in a reflective way — thanks to the culture of debate and the ability to critically analyze what they have learned during their studies. After all, it is one of our most noble tasks as university leaders to impart optimism to students, to teach them history, to demonstrate to them that there are alternatives to violence and populism, and, thus, to show them the way to a better future.”  Oliver Günther, Die diverse Universität (The Diverse University)

Q1. Your book Die diverse Universität (The Diverse University) opens with a question you say has occupied you since your student days in the early 1980s: why do societies invest in universities, and what do they get in return? That question felt abstract for decades — and then it became urgent. In the German interview you gave last year, you described how the Hamas attack on Israel, the war in Ukraine, and the rise of populism across Europe and the United States converged to make you realize that much of what seemed permanent — democracy, peace, academic freedom — can no longer be taken for granted. What shifted in your own thinking between the university president who held these questions as a longstanding personal interest and the one who decided the moment had come to write this book?

Oliver Günther: That is a very complex question but let me try to answer this succinctly. I – as many of your readers – grew up in a free democracy as a matter of course. There was freedom of speech, freedom of thought, and freedom of science. That has changed now. Hungary and Poland had populist right-wing governments for many years, with very negative effects on the democratic structures there. 

Russia invaded Ukraine, the conflict between Israel and Palestine seems to be going nowhere,  the US became a kleptrocracy. Who would have thought all of that even ten years ago?!  In such a situation, I find it extremely important for universities and their leaderships to rise and address these challenges. Universities are a crucial pillar of open democracies. That is why we have an obligation to defend the freedoms we have. And the best means to do that are our research, our teaching, and our writing.


Q2. You argue in your book that universities have inadvertently created a new kind of ivory tower — not the old one built on elitism and exclusivity, but one built on a narrowed public discourse in which topics like gendering and wokeness dominate academic debate while large parts of society feel left behind or actively alienated. That is a courageous self-criticism from a university president. How did you arrive at that conclusion, and what does it actually look like from inside a university leadership role when you sense that the institution you lead has drifted away from the society that funds it and depends on it?

Oliver Günther: I have always done my best to promote diversity and freedom of thought in the institutions I have been with. As a student at Berkeley I co-created a seminar “Social Implications of Computing”. Many of my professors and fellow students found this a waste of time at best. I am still very sympathetic to the agenda that is summarized by the term “woke” (well knowing that the definitions of this term now vary widely between different groups). But at the same time, we have to focus on our core mission as universities: excellent research, effective teaching, and active transfer. Transfer is important, in the sense of transfer into society. That means in particular that we should never tire to explain to the general public what we do and why we do it. Harvard is a prime example of a university that has failed utterly in this respect. We have to make sure that that doesn’t happen here in Europe. And the best way to do it is to focus on great research and teaching, as well as an effective exchange with the general public.


Q3. You are currently writing the English version of the book — which means bringing an argument originally grounded in the German university system to an international audience that includes American, British, and global readers, at a moment when academic freedom and university autonomy are facing particularly severe tests in the United States. What do you want the English version to say to an American academic leader or policymaker that the German version could not fully anticipate — and how has your thinking about the book’s central argument evolved as you have watched what has been happening across the Atlantic?

Oliver Günther: Well, I just mentioned Harvard. There has always been a lot of great research and teaching at Harvard. But the general American public got the impression that all they think about in Cambridge is unisex toilets and gendering. That is nonsense, of course, but it led the university into its biggest crisis ever. At the moment, the situation in the US is abysmal. Their current administration is about to destroy one of the biggest assets of their nation – a fantastic system of research universities. This will have long-term repercussions. We already see that the brightest young minds from Africa, South America, India, and the Arabic countries do not want to go to the US anymore for their Master or PhD. They prefer to go to Europe instead. They feel more welcome there, the quality of the research and teaching is essentially the same, and tuition fees are much lower than in the US and in the UK. (In Germany, actually, many great universities do not charge any tuition.) I certainly do not feel any Schadenfreude but this will certainly not help the US in the long term. In my book “The Diverse University,” I analyze the situation universities are in right now – both in the US and in Europe – and I make some suggestions how to change – always with the objective to contribute as much as possible to the common good.


Q4. You draw a direct line from scientific research to democratic stability — citing the COVID-19 vaccine developed out of a Deutsche Forschungsgemeinschaft project at Mainz as an example of how universities can deliver extraordinary public value under pressure, and in a remarkably short time. Yet that same pandemic also showed how quickly scientific authority can be contested, how easily misinformation spreads, and how difficult it is for universities to communicate complex, uncertain findings to a public that wants simple answers. What did the pandemic experience teach you about the communication gap between universities and society — and what has the university community failed to learn from it?

Oliver Günther: In recent years, universities have increasingly been in the public eye. A priori, that seems like a good thing. However, social media has also changed the rules of the game for scientific institutions.  For example, the threat posed by Covid-19 saw many politicians and regular citizens turning to science for advice. This was flattering for us scientists, but it also led to an increase in visibility and social responsibility that large sections of the scientific community were not used to. In the course of this debate, it became clear to many citizens that science does not always know everything. 

We may know that the earth is round and the moon is a satellite of the earth, but at the beginning of the pandemic, we did not know exactly what the coronavirus does to us humans and how it spreads. Of course, science cannot answer all questions with absolute certainty. Rather, there are many areas where there are majority and minority opinions, as well as ongoing discussions: namely, wherever research takes place, i.e., at the boundaries of proven knowledge. This is completely normal from a science point of view, but the public often misinterprets and misunderstands the contentious discourse inherent to science. We scientists must live up to this communicative challenge, as this is the only way to ensure that science can continue to be regarded as a credible authority. In retrospect, science has not done so badly, even in those difficult times of pandemic and mass hysteria. Quite the contrary: Science has done well. Nevertheless, it is now more important than ever to continuously communicate to the public what science and universities are all about: serving the common good. 


Q5. Generative AI is arriving in universities simultaneously from multiple directions — students using it for essays and assignments, researchers using it for literature review and data analysis, administrators exploring it for institutional processes, and national governments debating how to regulate it. From your position as president of a major research university, what is your honest assessment of how German and European universities are currently responding to GenAI — are they engaging with it seriously and courageously, or are they mostly reacting defensively and hoping the disruption passes?

And beyond the diagnosis: what concrete recommendations would you make to university leaders, policymakers, and faculty about the actions that need to be taken now to ensure a mindful, responsible, and educationally sound use of GenAI in higher education — one that harnesses its genuine potential without sacrificing the deeper purposes that universities exist to serve?

Oliver Günther: Our M.I.T. colleague Anant Agarwal said at a panel discussion in Kyoto in 2023: “AI in education is going to be very big. But at the universities I don’t see much happening yet.”[1]  Since then, we have seen some major changes in both public and private universities. The role of take-home papers and exams is increasingly being questioned. Admissions are based no longer on submitted essays alone. The tailoring of one’s teaching to specific students and the use of large language models, such as ChatGPT, has become quite common. More and more universities integrate a responsible use of AI into their teaching and examination routines. After all, AI is here to stay and our students need to understand how to use the available tools best to achieve their professional goals.

Already, adaptive learning empowers us to do more justice to the diversity of students.  AI-supported tools help to record the performance, talents, and motivation of individual students and to offer them personalized learning content based on this analysis. This is almost equivalent to a personal tutor for each student. Such targeted “tutoring” helps manage heterogeneity and will allow a higher proportion of students to succeed than would be the case with the usual one-size-fits-all teaching methodologies.

It is not unlikely that large language models will give rise to new cultural techniques and literacy skills that will need to be integrated into school and university curricula — techniques and skills that we need to teach our students to adequately prepare them for the challenges of the coming decades. It is safe to assume that, in the future, the majority of texts will be created by AI, at least in their initial version. Text creation will become a commodity, which naturally raises questions of intellectual property. Real-time voice-to-voice translation will also become broadly available shortly. For university teaching, it is important to not only recognize the availability of such powerful technologies, but also to help students use them efficiently and understand their limitations. With regard to large language models, this means training students in more than the formulation of queries (“prompt engineering”). We must teach them how to engage in a productive exchange with AI, including critical fact-checking.


Q6. There is a deep tension at the intersection of your book’s argument and the arrival of GenAI. Your book argues that universities must be open spaces that strengthen critical thinking, democratic participation, and the common good. GenAI can produce fluent, apparently well-reasoned text without genuine understanding, and can make it significantly easier for students to produce work that performs the appearance of learning without the substance of it. How do you think about GenAI’s impact on the thing universities are most fundamentally supposed to produce — not degrees, not publications, but citizens who can think independently and participate meaningfully in democratic life?

There is more to teaching than just imparting subject-specific content. Teaching means personality development. This also means confronting students with uncomfortable truths and with controversial views – including political views. In this respect, we can only agree with a report Yale published as early as 1974: “The history of intellectual growth and discovery clearly demonstrates the need for unfettered freedom, the right to think the unthinkable, discuss the unmentionable, and challenge the unchallengeable. […] [W]hoever deprives another of the right to state unpopular views necessarily also deprives others of the right to listen to those views.”


Q7. You describe universities as spaces where young adults develop not only professional knowledge but also as citizens of a free democracy — spaces that are essential for democratic society to reproduce itself across generations. If GenAI is increasingly capable of providing information, explanation, and even personalized instruction at scale and at low cost, what remains that only a physical university community — with its friction, its disagreements, its diversity of encounter — can provide? And is that remaining value sufficient to justify the scale of public investment that universities currently receive?

Digital tools, including of course GenAI, will remain an important element of academic education in the long term, and that is a good thing. However, face-to-face teaching and collaborative learning must remain at the forefront, in particular of undergraduate studies, as they still seem to be the best way to foster the personality development of our students. Only the physical presence on a diverse campus can provide the experience of controversial discourse and enrichment through exposure. Experiencing and debating with faculty and classmates from very different backgrounds than one’s own is an essential component of a modern education – an education that prepares our students for the complexities of our “new” world. With universities focusing on this kind of experience, tax money could hardly be better spent than on enabling many younger (and older) people to study there – at least for a short time. 


Q8. You envision universities as open spaces where controversial opinions can be heard, where red lines are identified and defended, and where society’s hardest problems can be addressed without fear or suppression. What does the university of the future actually look like, in your vision — not in ten years, but in thirty? How does it differ structurally, pedagogically, and in its relationship to society from the university of today — and what would need to change in governance, in funding models, and in the culture of academic institutions to make that future possible?

In Potsdam we are currently building a new campus for our Law School and our Faculty of Economics and Social Sciences. 1,000 faculty and staff, 6,000 students. It is a wonderful opportunity – but also a challenge – to build an infrastructure for the next 100 or more years of university life. The architecture is based on the principles I outlined above. We believe in physical presence. That’s why we need to make this space attractive. Students and faculty alike must enjoy to spend time there – inside and outside the classroom. While focusing on physical presence, the campus must also give plenty of opportunities to use digital media – alone or in groups. This inherent hybridity of academic life is maybe the biggest difference between today and the university we knew. In Potsdam, we will also have dormitories on campus – not as many as I would have liked but better than nothing. As we know from universities in many other countries, most notably the US and the UK, the resulting campus culture has an enormous impact on the personality development a good college wants to provide their students with.


Q9. Potsdam has a particular historical and symbolic weight — the city of Sanssouci, of Frederick the Great, of the Potsdam Conference. You quoted Voltaire in your interview — “I do not share your opinion, but I would die for your right to express it” — and noted that it feels more relevant today than ever. What does leading a university in Potsdam specifically mean to you in terms of historical responsibility — and how does that place shape how you think about the university’s role in the democratic moment we are living through?

When I look out of my office windows, I see the palace of Frederick the Great. My office building used to house his kitchen and housekeeping staff. So, of course, we think about those times frequently. Not everything Frederick did was indeed “great” but his openness to unorthodox ideas is a historical fact. In this vein, the University of Potsdam applies the principle “when in doubt, rule for tolerance.” This is also in keeping with the city’s history as an intellectual center of the European Enlightenment. Frederick’s friendship with Voltaire – which did not last – is also a fact. We celebrate this connection through our eponymous “Voltaire Prize for Tolerance, International Understanding, and Respect for Difference.”  It honors young researchers from anywhere in the world who are committed to the ideals of the Enlightenment, uphold them even in difficult situations, and oppose racism and discrimination. The aim is to support lecturers and students when it comes to expressing their opinions freely — both inside and outside the lecture hall. In difficult political times like ours, there can not be enough awards like this, I think.


Q10. You have spent decades as a computer scientist, a database researcher, a professor, and now as president of one of Germany’s most dynamic universities — and you are writing a book in German and translating it yourself into English about democracy, diversity, and the common good at a moment when all three are under pressure. That is not the career arc most computer scientists follow. What has the experience of leading a major public institution through some of the most turbulent years in recent European political history taught you about yourself — about what you value, what you are willing to defend publicly, and what kind of university president you wanted to be when you took the role versus the one you have actually become?

You say correctly “That is not the career arc most computer scientists follow.” That is correct, and maybe I was never the “typical” computer scientist. I was always interested in politics and on impact. When I was President of the German Informatics Society (GI), my motto was “More computer science into politics – More politics into computer science.” That was in 2011/12. So again, my kind of career is indeed somewhat atypical and it is not for everyone. But I would still wish for more computer scientists to branch out and – yes – BECOME politicians, managers, public administrators. My CS background is still helping me in my current job every day. We need that kind of technical background in public administration more than ever. And by the way: It was fun to change sides. I loved my work as a computer scientist but I love my job as an “administrator” just as much.

Qx. Anything else you wish to add?

I think we covered a lot of ground. Thank you for these great questions and the opportunity to speak out.

[1] URL: https://www.stsforum.org/news/the-summary-for-sts-forum-2023-20th-annual-meeting-is-now-available/, last accessed on: July 17th, 2026.

………………………………………………………………..

Oliver Günther has been President of the University of Potsdam since 2012. From 1993 to 2011 he was Professor of Information Systems at Humboldt University in Berlin, his research focusing on IT strategy, IT efficiency, and IT privacy/security.  He holds a Diploma in industrial engineering from the University of Karlsruhe and M.S. and Ph.D. degrees in computer science from UC Berkeley.

Aug 17 26

Why Enterprise AI Needs a Different Data Layer: David Flower on Agents, ACID, and the Honest State of Production AI

by Roberto V. Zicari

“ACID was, and will always be, the right answer to concurrent state access. Agentic AI just made ignoring it expensive instead of theoretical.”

Q1. When we last spoke in December 2017, you had just repositioned VoltDB around “translytics” — the combination of real-time transactions with real-time analytics — and you were seeing early adoption in fraud detection, telecoms, and mobile gaming. Eight years later, Volt Active Data is positioning itself squarely around AI infrastructure. Walk us through the honest version of that journey: what changed in the market, what changed in the product, and what stayed fundamentally constant about the core problem you are solving?

David Flower:  Eight years is a lifetime in this industry, so let’s be precise about the timeline. AI wasn’t a mainstream market disruptor until late 2022, when ChatGPT changed the conversation overnight. Back in 2017, translytics (the fusion of transactions and analytics) was itself early in both development and adoption. That trend hasn’t gone away; look at where every major data platform provider, including Databricks and Snowflake, is now heading. There has always been a desire to interact with data, whether in stream or table form, to extract value or act on it in real time. AI simply added another dimension to that same objective.

So we didn’t pivot to AI. The fundamentals we’re built on, which have always been about deciding correctly the instant a transaction happens, went from a niche requirement to a universal one. That shift made those fundamentals more necessary, not less relevant.

What we have done, however, is expanded our capabilities to meet market requirements. We added cross-data-center replication (XDCR), elastic scaling, zero-downtime upgrades, and now integration with AI through MCP-tool access, so agents can query us directly instead of hitting a stale replica.

In 2017, the problem was closing the gap between a transaction and its analysis. Today it’s closing the gap between an agent’s reasoning and an authoritative answer. It’s the same gap. It’s just that a lot more people have it now.

Q2. Volt has been positioning itself around AI, but you are a database company. Critics would say this is just another vendor jumping on the AI bandwagon — that the category label has changed but the product has not. Make the case against that criticism. What has genuinely changed in Volt Active Data’s architecture and capability since 2017 that makes it a different product for a different era, not just a rebranded one?

David Flower:  It’s a fair generalization to make about this market. We’ve all watched companies bolt the word “AI” onto their name or their elevator pitch. The difference is that we’re demonstrating AI’s value inside our product, not inside a slide deck. We built on our own foundations, but added new capabilities designed specifically for how agents behave differently from humans, and differently from the ML scoring workloads we served before.

Here’s the concrete version. The world has moved a long way since 2017, notably toward streaming and Kafka as the de facto way to move events around, rather than building complex native database applications to handle every transaction. Volt evolved with that shift, expanding from a pure transactional database into a data platform that ingests from many sources at once, including streams and traditional database clients. That makes us appropriate for deployments that need both working against the same dataset simultaneously.

That matters specifically for agents, because an agent reasons on whatever context it’s given, and the most current, real-time view of the world is typically whatever’s held in memory, which is exactly what Volt maintains. MCP is the protocol agents natively use to reach out to external systems like Volt and pull real-time context into their reasoning. Building clean, fast, structured access to live operational state for that purpose, and not just for storing data faster, is what’s genuinely new here.

Q3. You talk about ACID compliance as critical for AI agents. Most people associate ACID with traditional transactional databases, not AI infrastructure — and many of the most prominent AI data platforms deliberately trade off consistency for speed and scale. What is your reasoning for why ACID matters for agentic AI specifically, and what actually goes wrong in an AI agent workflow when the underlying data layer is eventually consistent rather than fully ACID-compliant?

David Flower:  First, it’s worth saying that “transactions” aren’t only financial. A transaction is any multi-step business operation that either fully happens or fully doesn’t. That idea applies just as much to a network capacity allocation as it does to a payment.

Take a non-financial example, such as an agent managing capacity for a network slice. The agent reads current available bandwidth, reasons about it, and recommends granting a new slice to a customer. If another process commits a competing allocation in the milliseconds between the agent’s read and the action executing, the agent’s recommendation was made against a fact that is no longer true by the time it matters. That’s a classic dirty read, and isolation (the “I” in ACID) exists specifically to prevent it.

Another shift is the read/write ratio. Early AI workloads were read-heavy: ask a model a question, get an answer back. Multi-agent systems are read-write. Agents are constantly writing traces, checkpoints, and intermediate decisions, often concurrently with other agents touching the same state. That reintroduces every classic concurrency problem distributed systems theory already solved.

And real business transactions are rarely as simple as “add to this, take from that.” They’re often many conditional steps forming a single logical operation. An agent reasons with whatever state it sees at the moment it looks. It doesn’t wait around to see how a complex transaction eventually resolves. So the state it sees must be meaningful. From the business’s perspective, a complex transaction either fully happened or it didn’t. A half-completed mess isn’t a valid state to reason from, which is exactly why it should be completed or rolled back, never left in between. ACID’s atomicity, consistency, and durability guarantees are what ensure an agent is always reasoning from a state that actually means something.

ACID was, and will always be, the right answer to concurrent state access. Agentic AI just made ignoring it expensive instead of theoretical.

Q4. Volt claims sub-10ms latency with full ACID compliance at scale. Google Spanner and Amazon Aurora are widely considered the benchmarks for distributed consistency at cloud scale, and CockroachDB and YugabyteDB are both well-funded and technically serious. Where does Volt genuinely lead, where does it genuinely trail, and what is the architectural reason for each — not the marketing answer, but the honest engineering one?

David Flower:  The honest starting point is that none of these systems are doing our job better or worse than us. They’re solving different problems, with different trade-offs made at different points in the stack.

Spanner’s own published numbers are roughly 10ms reads, and 50ms writes. That’s their number, not a knock. It’s the cost of TrueTime coordinating a globally consistent clock across continents, and it’s a genuine engineering achievement for that specific problem. We’re not trying to do what Spanner does. Spanner, CockroachDB, YugabyteDB, and Aurora all pay for their scale and geographic reach by trading away some degree of consistency. Each makes a slightly different trade, but none of them operates at the same strict ACID level Volt does.

Volt pays for that consistency at design time instead of at query time. You declare partition keys up front and express transactions as deterministic stored procedures. That changes what a latency comparison even means. Comparing a 1–2ms round trip in Volt to a several-millisecond round trip elsewhere isn’t comparing like for like. In Volt, the entire business transaction executes as one unit, server-side, inside that 1–2ms. In the other systems, a single round trip is often just a fragment of the logic (one read, one write, one loop iteration), so accomplishing the same business transaction usually takes several round trips strung together on the client side, not one.

Where we clearly lead beyond raw latency is tail predictability. No locks, no latches, and fully deterministic execution mean no long-tail latency spikes, meaning results stay predictable at the 99th percentile, not just on average. Our active-everywhere replication (XDCR) is also more predictable from a transaction-latency standpoint than the multi-region approaches the others use.

If we’re going to be honest about where we trail, it’s multi-partition transactions. Volt asks you to architect your data and transactions to run against a single partition wherever possible. That constraint doesn’t fit every business use case cleanly. That’s a genuine trade-off. Our lane is bounded. It’s high-value operational state — balances, sessions, entitlements, network capacity — where being wrong has a real cost. We don’t try to win on unbounded data volume or general-purpose SQL surface area. For those needs, one of the other three is probably the right call.

Q5. In 2017 you described VoltDB’s sweet spot as applications where “the window of opportunity is available in the fast data stream process and once passed the opportunity value diminishes.” That description fits fraud detection perfectly. How well does it fit the emerging agentic AI use cases you are now pursuing — and are there categories of AI agent application where that real-time decisioning model does not actually apply, and where a different architectural approach would serve better?

David Flower:  That philosophy has never been more important, and it now reaches well beyond fraud or financial transactions. It’s become the basis of value extraction from real-time data generally, alongside the validation and accuracy of the decision or recommendation itself.

It still holds true for most of what has rapidly become the agentic market. However, there’s a growing slice where it genuinely doesn’t. It fits fraud, charging, and entitlement checks, where a late “allow” is a wrong “allow” regardless of how smart the thing that produced the recommendation was. It doesn’t fit a research agent spending an hour synthesizing an analysis, a planning agent working through a multi-step remediation, or overnight document reconciliation. Nothing there closes in milliseconds. Being a few minutes late costs nothing.

The deeper point is less about speed and more about validity. What matters most is ensuring the data an agent reasons from is as current and consistent as it can possibly be. Reasoning from a half-executed transaction (a state that isn’t valid from the business’s perspective) is a bad outcome whether the agent takes a millisecond or an hour to act on it.

Q6. In 2017 you told me that most enterprise AI and ML was still post-transaction — analytics on the back end rather than decisions in the stream. That was eight years ago. Today, there is enormous pressure on enterprises to move AI inference into the critical path of real-time operations. What is your honest assessment of where most enterprises actually are in that journey in 2026 — and what is the biggest gap between where they think they are and where they actually are?

David Flower:  I’d strongly agree that every company is facing enormous pressure to move AI inference into the critical path. But those decisions and recommendations must be accurate. There is no value in being fast but wrong, and for a lot of these use cases, being fast but wrong could be genuinely catastrophic for the business.

Most “AI in production” today is still the 2017 pattern with a much smarter model attached to it: insight gets generated, and a person or a batch process still has to act on it. The gap between where enterprises think they are and where they actually are comes down to context and determinism. In other words, whether the outcome an AI system produces can actually be trusted.

Concretely: pilots are deliberately run on clean, curated data. Production data is live and messy. MIT’s 2025 “GenAI Divide” research found that the overwhelming majority of generative AI pilots fail to show measurable results once they move to production. The consistent finding across that research and others like it is that it’s rarely the model. When accuracy drops after go-live, teams tend to blame the model. Usually the model is fine. This is the first time we’re reasoning against stale or inconsistent state, because production data was never as clean as the pilot dataset it was built and tested against.

That’s the gap: enterprises think they have an AI maturity problem. Underneath it, most of the time, they have a data problem that existed before the AI project ever started, and was invisible right up until an agent began acting on it continuously instead of a person glancing at it once a day. To be fair, we have also seen a real, encouraging uptick in enterprises adopting ML models specifically for anomaly detection, and that’s one place the maturity curve is genuinely moving in the right direction.

Q7. Most research still shows that the majority of enterprise AI deployments are in pilot mode rather than production. You are selling infrastructure for production AI. What do you actually hear from enterprises about why pilots stall — and what does that tell you about what the market genuinely needs that it is not yet getting from the AI infrastructure category as a whole?

David Flower:  What we generally hear isn’t “the data layer failed” in those words. It’s “the model got worse in production” or “we can’t explain a decision after the fact.” Those are the same root cause stated differently.

I’d add that fear is a major factor too, particularly for mission-critical, production environments. A good example is autonomous networks in the mobile CSP market. Every operator is chasing the silver bullet, but very few are actually ready to let AI run the most critical assets of their business unsupervised. That’s where proven performance fundamentals and evidence-based validation become essential. Trusted AI outcomes need a deterministic foundation underneath them, not just a confident-sounding model.

There are plenty of numbers to back this up. Deloitte’s 2025 research found only 11% of enterprises have agentic AI actively running in production, and MIT’s 2025 research found the overwhelming majority of generative AI pilots never show measurable results. That’s a significant gap given the level of investment in the space.

We see the trust question play out directly with customers. In telecom, we’ve seen customer care agents get replaced by AI specifically because the outcome is auditable. A customer complains their calls always drop driving home past a certain location; the AI agent reviews the network data and call records, determines the pattern is real, applies a goodwill credit (and can see and audit exactly how often this has happened for that subscriber), and triggers a network engineering ticket to review capacity at that location. It’s contained, and every step of it is auditable. Now compare that to letting the same agent autonomously go fix the network itself, allowing it to add capacity, spend on additional trunk routes and interconnects, with no human in the loop. The cost to the operator, and the stakes if it gets something wrong, get a lot scarier very quickly. AI has to be auditable, and it has to earn trust before it’s let loose on more critical, more autonomous tasks. That’s exactly the gap the AI infrastructure category as a whole isn’t yet closing: plenty of tooling for storing and orchestrating agents, very little built for agent-paced, auditable decision authority on live operational data.

Q8. You work with partners across the data ecosystem — streaming platforms, AI frameworks, cloud providers. The honest question is: does Volt fit alongside other data platforms as a complementary layer, or is your strategic goal ultimately to consolidate and replace them? And if the answer is “complement,” where exactly is the boundary — what does Volt do that you would tell a customer not to try to do with Kafka, Flink, Snowflake, or a vector database?

David Flower:  We absolutely act as a complementary layer, and the boundary is concrete, not a vague “we play nicely together” answer.

Don’t ask Kafka or Flink to own a decision. They move and process events extremely well, but they were never built to be the authority, and they don’t offer a full decision audit trail on their own. Don’t ask Snowflake “what should happen right now.” It answers “what happened” and “what’s the pattern.” It’s worth noting we’ve seen several major data lake and warehouse vendors investing in, acquiring, or exploring real-time transactional technology recently. That’s a clear sign they recognize the same gap we’ve been pointing at for years. And don’t ask a vector database to decide anything: similarity search finds a plausible match; it doesn’t apply business logic.

A good, live example of the complementary model is our partnership with Ocient. Volt holds live operational state, such as current balance, active session, and real-time network state. Ocient holds petabyte-scale historical pattern data. An agent reasoning from both can say, in effect, “here’s what’s happening right now, in the context of what’s historically normal.” The result is a meaningfully better answer than either platform could produce alone.

Q9. You have been running Volt Active Data — including through a rebrand from VoltDB — for nearly a decade. That is a long time to lead a company in a market that has changed as dramatically as data infrastructure has changed. What has been the hardest strategic decision you have had to make in that period — the one where you were most uncertain, where the data did not give you a clear answer, and where you had to lead from conviction rather than evidence?

David Flower:  First, we haven’t changed our core DNA or the class of problems we solve. Volt was created by one of the world’s smartest technical minds in this field, Mike Stonebraker, specifically to solve complex, mission-critical problems. The strength of our global enterprise customer base is the evidence that the thesis was right.

Given that, the hardest decision hasn’t been what to build. It’s been when to make a strategic shift. AI is the obvious recent example: it was hyped for years, but with not much substance underneath it in the early days. Move too early on a wave like that, and you waste a huge amount of valuable capital chasing noise. Move too late, and you miss the boat entirely. On balance, I think we made the right calls at the right time, and we’re seeing that pay off now. But we have to stay agile, because you genuinely never know what the next wave is going to be.

The judgment call underneath all of it is telling the difference between noise and real business value before the market has made it obvious. AI and streaming both turned out to be real. Plenty of other trends haven’t been. The metaverse comes to mind. It never ended up being particularly relevant for what we do. Knowing which one you’re looking at, before the evidence is in, is the actual hard part.

If I had to point to a cultural anchor through all of it: one of our earliest customers told us, “it just works.” We’ve been told since that we’re so reliable we’re almost boring. Our best customers, such as the fraud and billing teams who can’t afford to be wrong, are the ones who keep pulling us back to that narrow, deterministic story whenever the market pressure says we should be broader. That pull has been more useful to me than any market data.

Q10. Looking at the database landscape in five years, given the trajectory of agentic AI — where every agent needs safe, fast, consistent access to operational data at machine speed — where does Volt Active Data sit in that picture, and what would have to be true for that future to play out in your favor rather than in the favor of the cloud hyperscalers or the next wave of purpose-built AI data platforms?

David Flower:  I’ll be honest. This question feels like it was written for us to show off our own positioning, so let me try not to overdo it.

The landscape will split further before it consolidates. The operational tier, the analytical tier, vector stores, and artifact/trace stores are genuinely different problems. That’s not one problem that collapses into a single vendor, no matter how good that vendor is. What I do expect to matter more everywhere is atomic decision recording. Regulators are moving toward mandatory explainability for automated decisions, which means the same discipline that’s been required in regulated industries for fifteen years is going to be expected far more broadly as agents take on more consequential work.

In answer to the “what has to be true for us” question, it’s that enterprises need to keep hitting the wall where “good enough” consistency breaks the moment it’s an agent, not a person, reading and acting on the data. Hyperscaler bundled options will keep being the right, cheaper choice for a lot of workloads, and we don’t need to win that fight. We need to keep being the correct answer where correctness is genuinely non-negotiable, and that slice of the market is growing, not shrinking.

Qx Anything you wish to add?

David Flower:  Just that none of this is new to us. Fraud, telco, and network teams have needed correct, immediate, explainable decisions for the fifteen years we’ve been around. Agentic AI didn’t invent that requirement, but it did expose it to a much bigger audience and made the cost of getting it wrong visible overnight.

Resources

Facing the Challenges of Real-Time Analytics. Interview with David Flower, ODBMS Industry Watch, December 19, 2017

………………………..……………………..

David Flower, President & CEO, Volt Active Data

David Flower brings more than 30 years of experience within the IT industry to the role of President and CEO of Volt Active Data. David has a track record of building significant shareholder value across multiple software sectors on a global scale through the development and execution of focused strategic plans, organizational development, and product leadership.

………………….

Follow us on X

Follow us on LinkedIn

Jul 27 26

A New Era for MySQL: Heather VanCura and Jason Wilcox on Open Source, Community Governance, and Where MySQL Is Headed

by Roberto V. Zicari

“Through transparent roadmaps, community-driven collaboration, contributor programs, and the MySQL Governance model, we aim to create an environment where innovation can accelerate while preserving the reliability, compatibility, security, and operational excellence that organizations around the world depend on.”

Q1. Oracle has announced a “new era” of MySQL community engagement at MySQL’s 30th anniversary. Can you walk us through what specifically prompted this strategic shift, and what concrete changes can the community expect to see in how Oracle approaches MySQL development and governance?

HVC: Throughout 2025 we celebrated 30 years of MySQL and reflected on the past and present, but more importantly, the future. The MySQL Community team sought feedback from around the globe on how to lead the next generation of MySQL innovation and open source collaboration. We came to Jason in November and shared that feedback and proposed a plan to rebuild community trust. By December we agreed on a plan, calling it a new era of Community Engagement.

We have entered a deeper collaboration with the MySQL Community, focused on faster innovation, greater transparency, deeper community collaboration, and expanding the ecosystem. Starting with the April 2026 release, we’re delivering more features directly into the MySQL Community Edition core while preserving the stability customers rely on.

As part of this effort, we introduced the MySQL Governance model, which provides clear pathways for participation, community leadership, and long-term collaboration with the broader MySQL ecosystem. Together, these initiatives are designed to build deeper trust, accelerate innovation, and grow the MySQL ecosystem.

Q2. One of the most significant announcements is moving previously commercial-only features into the MySQL Community Edition. What drove this decision, and what other enterprise features are you planning to bring to the community edition in the coming months?

JW: Both our Community and our Customers are asking for stability and faster innovation. At the same time, they want more visibility into our roadmap and a stronger voice in shaping it. Driving some previously Enterprise-only features into the Community Edition addresses both needs, while we not only deliver those features, but we also build, prioritize, and deliver new features and innovations into MySQL.

With the GA of MySQL 9.7.0 LTS, MySQL moves from the 9.x innovation series to a new Long-Term Support release line. This begins the 9.7.x LTS series, giving users a stable branch to standardize on while continuing to build on the innovation delivered through the 9.x cycle.

This release matters not only because it establishes the next LTS baseline, but because it reflects a broader direction for MySQL. Over the last several releases, we have talked about giving users earlier visibility into what is coming, broadening access to important capabilities, and working more openly with the MySQL community. With MySQL 9.7.0 LTS, that direction is reflected in the product itself.

Several capabilities previously limited to MySQL Enterprise Edition are now available in MySQL Community Edition, while Dynamic Data Masking is now available in MySQL Enterprise Edition. Together, these changes make MySQL 9.7.0 LTS a meaningful release for DBAs, developers, and operators across both editions.

More capability in MySQL Community Edition

One of the biggest themes in MySQL 9.7.0 LTS is the continued expansion of MySQL Community Edition. Across 4 major technical areas, this release delivers 8 notable new Community Edition capabilities — a substantial broadening of what DBAs and developers can do with Community Edition.

The 4 major areas

  • Replication observability and HA behavior
    • Multi-threaded applier extended statistics
    • Automatic Eviction & Rejoin
    • Up-to-date Aware Primary Election
  • Telemetry and observability integration
    • Telemetry / OpenTelemetry support
  • Modern application development
    • MySQL JSON Duality Views
  • Query optimization and performance
    • Hypergraph Optimizer
    • Profile-Guided Optimization (PGO)

Q3. Some community members have expressed concerns about MySQL’s development velocity and commit rates. Jason, as SVP of Data Services, what specific steps are you taking to address these concerns, and how do you plan to balance cloud service development with core MySQL innovation?

JW: Oracle has invested heavily in MySQL since 2010, and we hear feedback from the community. People want to see that investment show up in a more visible way, especially through faster delivery in the open. We’re working on that in a few concrete ways: getting more features into MySQL Community Edition, sharing more of the roadmap and worklogs, using Early Access releases to get feedback earlier, and creating more public forums where contributors can talk directly with the MySQL engineering team.

We’re also putting more structure around how people can participate, through the MySQL Governance model, contributor summits, design discussions, and clearer contribution paths. The goal is straightforward: be more open about where MySQL is going and give the community more practical ways to influence priorities, test features earlier, report issues, and contribute improvements. Cloud and core MySQL are not separate priorities for us — the core database is the foundation for Community, Enterprise, and HeatWave, so continued innovation in MySQL itself remains central to everything we’re doing.

In addition to accelerating innovation, we are creating more opportunities for community participation through public roadmaps, Early Access releases, public discussions, contributor summits, and the MySQL Governance model. Together, these initiatives provide greater transparency into our priorities while creating structured mechanisms for contributors to participate, provide feedback, and help influence the future direction of MySQL.

Q4. Oracle has published the MySQL Community roadmap and promised to facilitate community contributions through worklogs and bug reports. How will this differ from past practices, and what mechanisms are you putting in place to ensure transparent, bidirectional communication between Oracle’s engineering team and external contributors?

HV: With our Community Engagement Plans, MySQL customers and users get the best of both worlds: enterprise-grade stability and faster access to innovation. They’ll also have greater visibility into what’s coming and more opportunities to provide input, which helps them align MySQL with their own technology roadmaps. In addition to publishing select worklogs and CVE information, we have continued Labs for new features and early access releases leading up to the 9.7 launch, which will continue in future releases, with our next Early Access planned for early July. These provide valuable insight and transparency to community members and invaluable feedback to the engineering team. 

We have organized a series of public discussions (four so far), with a fifth planned for July, as well as established a quarterly Contributor Summit and regular design meetings under the MySQL Governance model. The first Contributor Summit took place in May 2026, with a design meeting held the week prior. The next Contributor Summit is scheduled for August 2026 in Broomfield, Colorado.

The governance model provides structured pathways for participation through code contributions, testing, documentation, reviews, technical discussions, and community leadership. It introduces clearly defined roles—including Contributors, Committers, Project Leads, Core Project Leads, a Steering Committee, and a Vulnerability Group—to help ensure transparent collaboration while maintaining MySQL’s standards for quality, stability, compatibility, and security.

In the last quarter, we also published the MySQL Developer Guide, which describes how to effectively contribute and participate in the evolution of MySQL.

To catch up on previous discussions, see highlights from earlier sessions:

Q5. How do you see the MySQL governance structure evolving to give the community a stronger voice while maintaining Oracle’s stewardship?

HV: The MySQL Governance model is a key part of how we are evolving community participation while maintaining Oracle’s long-term stewardship of the project. The model is built on principles of transparent processes, merit-based participation, shared stewardship, and a commitment to quality, stability, compatibility, and security.

Oracle remains the primary steward of MySQL while creating clearer pathways for the community to participate in shaping the project’s future. Community members can contribute through code, testing, documentation, bug reports, design discussions, and reviews. As contributors gain experience and demonstrate sustained engagement, they can take on greater responsibilities through defined governance roles.

The model also introduces a Steering Committee that brings together perspectives from Oracle, users, customers, hyperscalers, and the broader open source ecosystem to help guide long-term priorities, governance evolution, ecosystem growth, and community engagement.

Together with public roadmaps, Early Access releases, GitHub collaboration, contributor summits, and design meetings, the governance model creates a structured framework for community participation while preserving the engineering excellence and operational stability that organizations around the world depend on.

Q6. PostgreSQL has been gaining ground with features like pgvector for AI workloads, while MySQL faced criticism for lack of similar capabilities. How does Oracle plan to ensure MySQL remains competitive not just with PostgreSQL, but also with cloud-native databases and newer entrants in the database market?

JW: MySQL offers a uniquely predictable and stable operational model at global scale, combined with strong performance and ease of use. Backed by Oracle, it delivers enterprise-grade reliability while maintaining the flexibility and innovation of open source. We will continue to collaborate with the community to deliver innovations based on our published roadmap into MySQL Community Edition. 

Q7. The move of MySQL into Oracle’s cloud organization raised concerns about resource allocation. Can you address these concerns and explain how Oracle is ensuring MySQL has the engineering resources it needs to execute on this new community-focused vision?

JW: MySQL’s success has always come from the combination of strong stewardship and a vibrant community. Oracle continues to invest deeply in both. What’s new is increased transparency, stronger engagement with the community, and more structured ways for contributors, partners, customers, and ecosystem participants to help shape MySQL’s future through the MySQL Governance model and related community programs.

These investments complement our continued engineering investment in MySQL Community Edition, MySQL Enterprise Edition, and MySQL HeatWave.

Q8. You’ve mentioned expanding collaboration with Linux distributions, particularly Canonical and Ubuntu, as well as supporting major open source projects like WordPress and Drupal. What does this ecosystem support look like in practice, and how will Oracle work with companies that some might consider competitors in the MySQL space?

HV: We have built relationships and communication between the MySQL Community Team and open source maintainers to ensure the pathways are smooth for projects to build their projects and platforms using MySQL.  We continue to strengthen communications and remove barriers to collaboration. 

That spirit of collaboration is reflected in the MySQL Governance model and community engagement efforts. The recent Contributor Summit brought together Oracle engineers and contributors from organizations including Amazon, Google, Percona, ProxySQL, Readyset, VillageSQL, and participants from across the broader MySQL ecosystem, including MariaDB, to share ideas and help shape the future of MySQL.

We continue to focus on growing and expanding the MySQL ecosystem, referencing the analogy of a rising tide lifting all boats. Growing the community and bringing more collaboration and alignment makes us all stronger together and creates opportunities throughout the ecosystem.

Q9. For organizations currently running MySQL in production, what’s your message about long-term support and the roadmap? With MySQL 8.0 approaching end of life and MySQL 9.7 LTS published in April 2026, how should enterprises plan their migration strategies and what assurances can you provide about stability and backward compatibility?

JW: MySQL offers a uniquely predictable and stable operational model at global scale, combined with strong performance and ease of use. Backed by Oracle, it delivers enterprise-grade reliability while maintaining the flexibility and innovation of open source. 

The release introduces a new long-term support version of MySQL Community Edition, and MySQL Enterprise Edition, along with expanded feature delivery into the core, early access capabilities, and the first phase of our enhanced transparency and community engagement model.

Q10. Looking beyond the immediate announcements, what is Oracle’s five-year vision for MySQL? How do you see MySQL evolving to meet the demands of AI workloads, cloud-native architectures, and modern developer expectations while preserving the simplicity and reliability that made it the world’s most popular open source database?

JW: MySQL powers everything from startups to hyperscale platforms. It’s used by companies like Uber and Booking, and underpins major platforms like WordPress and Ubuntu. That breadth of adoption is a strong validation of its reliability and scalability.

 The vision is simple: build MySQL in the open with the community, accelerate innovation without sacrificing quality or stability, and continue to scale and grow the ecosystem around the world’s most widely used open source database platform.

A key part of that vision is establishing a sustainable governance framework that enables broader participation, develops future community leaders, and creates stronger connections between Oracle, contributors, customers, partners, hyperscalers, and the broader open source ecosystem.

Through transparent roadmaps, community-driven collaboration, contributor programs, and the MySQL Governance model, we aim to create an environment where innovation can accelerate while preserving the reliability, compatibility, security, and operational excellence that organizations around the world depend on.


Jason Wilcox
Senior Vice President, Data and AI Platform, Oracle Cloud Infrastructure (OCI)
Jason Wilcox leads the Data and AI Platform organization at Oracle Cloud Infrastructure (OCI), overseeing the design and development of OCI’s data platforms, AI infrastructure and platform services, and open source technologies. His portfolio spans cloud-scale data services, data processing and integration platforms, operational services for AI workloads, and widely adopted open source technologies that developers and enterprises rely on to build modern applications. These services help customers manage and use data, run AI workloads, and operate secure, reliable, and scalable systems on OCI.

Heather Vancura
Vice President, External Standards & Community Engagement, Oracle Cloud Infrastructure (OCI) Heather VanCura is Vice President of External Standards & Community Engagement at Oracle, where she leads Java Community programs and the MySQL Community Outreach team. With over 20 years of experience at Oracle and Sun Microsystems, she is a central figure in the global ecosystem, focusing on community growth, engagement, and standardization efforts.

………………….

Follow us on X

Follow us on LinkedIn

Jun 30 26

What I Didn’t Learn in Medical School: Mathias Goyen on AI, Judgment, and the Human Side of Healing

by Roberto V. Zicari

“When patients say that AI listens better than their doctor, they are rarely making a statement about empathy. They are making a statement about time.”

Q1. Your book (*) argues that medical schools teach the technical anatomy of disease but not the anatomy of human hopes and fears, that physicians learn to diagnose but not always to truly listen. As AI systems increasingly match or exceed physicians on the technical and encyclopedic dimensions of medicine, you suggest the physician’s role as a trusted human ally becomes more important, not less. But that humanistic competency is, by your own account, something doctors are largely left to acquire on their own, often through harsh and humbling experience.

What would it actually take to teach this deliberately, and do you believe medical education as an institution is capable of changing fast enough to do so before an entire generation of physicians has already been shaped by the system as it exists today?

Mathias Goyen: The arrival of AI has created a fascinating paradox. The more capable our technology becomes at processing information, the more valuable distinctly human capabilities become. For centuries, medicine has largely defined excellence through knowledge, diagnostic accuracy, and technical skill. Those qualities will always remain essential, yet they are no longer sufficient on their own because information has become increasingly accessible while judgment, trust, and the ability to accompany another human being through uncertainty have become the true scarce resources.

This is precisely where I believe medical education faces its greatest challenge. We still devote enormous energy to teaching students how diseases behave, yet comparatively little attention is given to how people behave when they become patients. We teach physiology, pathology, pharmacology, and anatomy with remarkable rigor, but far less time is spent understanding fear, uncertainty, hope, grief, or the psychological complexity that accompanies almost every serious diagnosis. Those aspects are often treated as something physicians will simply acquire through experience, as though compassion naturally emerges after enough years on the wards. Experience certainly matters, but experience alone is an unreliable teacher. Some physicians become wiser through it, while others simply become more efficient.

The encouraging news is that I do not believe these qualities are beyond teaching. What we cannot teach through lectures alone can be cultivated through deliberate exposure to complexity. Medical students should spend more time observing difficult conversations than memorizing another list of rare syndromes. They should regularly reflect on situations in which there was no perfect answer. They should discuss uncertainty with senior clinicians who are willing to admit that medicine is often practiced without complete certainty and that wisdom frequently consists of choosing responsibly among imperfect alternatives rather than identifying a single correct solution. In other professions, including aviation, the military, and executive leadership, reflection after difficult situations is considered an essential part of professional development. Medicine still tends to reward certainty even when uncertainty is the daily reality.

AI makes this transformation more urgent, not because it diminishes physicians, but because it changes where physicians create value. If machines increasingly become exceptional at organizing knowledge, physicians must become exceptional at helping people navigate that knowledge. Patients will continue to need someone who can interpret information within the context of an individual life, balance competing priorities, communicate honestly when certainty is impossible, and remain present when medicine reaches its limits. None of these responsibilities become less important because an algorithm is available. If anything, they become more central to the profession than ever before.

Can medical education change quickly enough? I believe it can, but only if we accept that the future physician requires a broader definition of competence than the one that has guided us for generations. Medical schools have repeatedly demonstrated their ability to adapt when science demanded it. We incorporated molecular biology, genomics, and digital medicine into our curricula because they became indispensable. We should now show the same determination in teaching judgment, communication, ethical reasoning, adaptability, and the ability to lead patients through uncertainty. These are not soft skills that merely complement medical expertise. In the age of AI, they increasingly define it.

Q2. Patients increasingly arrive at appointments having already consulted ChatGPT, Gemini, or Claude, sometimes reporting that “AI listens as no doctor did before.” That is a remarkable and uncomfortable statement about the current state of clinical encounters.

What does it actually mean, in practice, for a physician to compete with, or more usefully, to integrate, a technology that can offer patients more time, more patience, and the appearance of being heard, within a healthcare system that structurally cannot offer physicians the time to do the same?

Mathias Goyen: I actually do not believe physicians should think of themselves as competing with AI. The moment we begin framing the relationship as a competition, we have already misunderstood what patients are really telling us.

When patients say that AI listens better than their doctor, they are rarely making a statement about empathy. They are making a statement about time. AI never interrupts. It never looks at the clock. It never appears distracted by the next patient waiting outside the door. It allows people to finish their thoughts before responding. That experience alone can create a powerful feeling of being heard, even though the technology itself experiences neither compassion nor understanding.

That observation should not make physicians defensive. It should make us reflective. It forces us to ask a difficult question about our healthcare systems rather than about our technology. Have we gradually created an environment in which efficiency has become so dominant that patients increasingly value uninterrupted attention as much as medical expertise? I suspect the answer is yes.

Ironically, I see this development as an opportunity rather than a threat. If patients arrive having already explored their symptoms with AI, the consultation no longer needs to begin with the simple transfer of information. Instead, it can begin at a much more meaningful level. The physician can help patients interpret what they have learned, distinguish probable explanations from unlikely ones, place isolated facts into the context of an individual life, and openly discuss uncertainty where uncertainty genuinely exists. In other words, the conversation can move away from information retrieval and toward judgment.

This also changes the physician’s role in a subtle but important way. Historically, physicians often served as the primary source of medical knowledge. Increasingly, they become trusted interpreters of knowledge that is already available to everyone. I consider that an evolution rather than a loss. Trust has never depended simply on possessing information that others do not have. Trust emerges when someone helps us understand what information actually means for our own lives.

The real danger, therefore, is not that AI becomes too patient. The danger is that healthcare systems conclude that because AI can provide unlimited conversational capacity, human conversation becomes less necessary. That would fundamentally misunderstand why patients seek physicians in the first place. Patients do not simply come to receive answers. They come to share responsibility for decisions that may profoundly affect their lives. They want someone who can recognize when uncertainty remains, explain why different options carry different consequences, and occasionally say, “I do not know, but we will work through this together.” Those moments create trust in a way that no technology can replicate.

Ultimately, I hope AI will not replace the conversation between physician and patient but elevate it. If AI can assume much of the informational and administrative workload, physicians should have greater freedom to focus on the conversations that require wisdom rather than recall, presence rather than speed, and judgment rather than computation. Whether that vision becomes reality, however, depends far less on the technology itself than on the choices healthcare organizations make about how the time that AI creates is ultimately used.

Q3. You write that the physician today carries not only a stethoscope but also data, and that patients now often have access to the same data and the same AI tools that physicians have. This symmetry is historically unprecedented in medicine. For healthcare leaders and policymakers far beyond any single organization or technology vendor, what do you believe is the single most important structural or educational change needed to ensure that this newly symmetrical relationship between doctor and patient becomes a source of trust and collaboration, rather than confusion, distrust, or a further erosion of the time available for genuine human connection?

Mathias Goyen: If I had to identify a single priority, it would not be the introduction of more technology. It would be the deliberate cultivation of judgment as a shared competency between physicians and patients.

For most of medical history, knowledge itself created an asymmetry. Physicians possessed information that patients simply could not access. Today that asymmetry is rapidly disappearing. A patient can read scientific publications, access clinical guidelines, ask sophisticated questions of large language models, and arrive in the consultation having accumulated an extraordinary amount of information. That development should not be feared. An informed patient is not a threat to medicine. Quite the opposite. Mutual understanding has always been the foundation of shared decision making. The challenge is that access to information is not the same as the ability to interpret it. Modern medicine generates probabilities rather than certainties. Imaging findings require clinical context. Laboratory values depend on medical history. Risk predictions require value judgments about what matters most to an individual patient. AI can organize remarkable amounts of information, yet deciding what deserves attention, what uncertainty remains acceptable, and which option best reflects a person’s preferences continues to require human judgment.

This is why I believe healthcare systems should move beyond thinking primarily about digital literacy and begin focusing much more deliberately on decision literacy. Physicians need to become better at explaining uncertainty without undermining confidence. Patients need to become more comfortable understanding that medicine rarely offers absolute answers and that reasonable experts can occasionally reach different conclusions while acting in good faith. Trust grows when uncertainty is acknowledged honestly rather than hidden behind artificial certainty.

There is also an important leadership responsibility. Healthcare organizations should resist the temptation to measure success exclusively through productivity metrics, waiting times, or numbers of consultations completed. Those indicators matter, but they tell us remarkably little about whether patients actually understood the decisions that were made together. If AI allows us to create more informed patients but simultaneously leaves physicians with even less time to interpret information collaboratively, we will have solved the wrong problem.

Ultimately, I believe the relationship between physicians and patients is becoming less hierarchical and more collaborative. That is one of the most profound cultural shifts medicine has experienced in generations. The physician’s authority will increasingly arise less from exclusive access to knowledge than from the ability to guide thoughtful decisions in situations where knowledge alone is insufficient. Patients, in turn, become active participants rather than passive recipients of care. I consider that an extraordinarily positive development because trust built through partnership is ultimately stronger than trust built through dependency.

If we succeed in creating healthcare systems that value dialogue as much as diagnosis, AI may become one of the greatest opportunities medicine has seen in decades. If we fail, we may discover that we improved the flow of information while unintentionally weakening the relationships that give information its meaning.

Q4. You made the unusual transition from practicing physician and academic radiologist to global Chief Medical Officer inside a major commercial MedTech company. That move places you in a position that carries an inherent tension: the Hippocratic commitment to the patient’s interest above all else, and the commercial reality of an organization whose technologies must ultimately sell and generate returns for shareholders.

How do you personally navigate that tension in your day-to-day decisions, and what would you say to a young physician who is skeptical that genuine humanistic medicine and a senior leadership role inside a commercial healthcare company can coexist with integrity?

Mathias Goyen: This question assumes a tension that certainly exists, yet perhaps not in quite the way many people imagine. Throughout my years in industry, I rarely experienced the debate as one between doing what is right for patients and doing what is right for the business. More often, I experienced it as a question of time horizon.

Healthcare is unusual because trust accumulates slowly and can disappear remarkably quickly. Physicians recommend technologies because they believe they improve patient care. Hospitals invest because they expect meaningful clinical value over many years. Companies build reputations over decades rather than quarters. When viewed from that perspective, serving patients well and building a successful business are not competing objectives. They are deeply interconnected. Commercial success becomes sustainable only when clinicians genuinely believe that a company helps them care for patients more effectively.

As a physician working inside industry, I always considered my role somewhat different from many other leadership positions. I was not there to replace commercial thinking. I was there to complement it with clinical perspective. Every discussion about product development, workflow, AI, education, or implementation eventually led me back to the same questions. Does this solve a meaningful problem? Will it make a physician’s work easier, more thoughtful, or safer? Will it ultimately improve the patient’s experience or outcome? Those questions do not eliminate difficult business decisions, but they provide a remarkably consistent compass.

I also learned something that surprised me. Before joining industry, I imagined companies primarily as organizations that develop technologies. Over time I came to realize that their greatest influence often lies elsewhere. They shape education. They convene experts from around the world. They invest in research. They help translate scientific discoveries into everyday clinical practice. They create ecosystems that individual hospitals or universities could rarely establish on their own. When those responsibilities are approached thoughtfully, industry becomes an important partner in advancing healthcare rather than merely supplying it.

To a young physician who is skeptical about entering industry, I would simply say that integrity does not depend on the logo on your business card. It depends on whether you remain intellectually honest about whom your decisions ultimately serve. Good people can make poor decisions in universities, hospitals, governments, or companies. Equally, principled leadership can exist in all of those environments. The ethical responsibility travels with the individual rather than the institution.

Perhaps my years in industry strengthened rather than weakened my belief in humanistic medicine. They allowed me to appreciate healthcare from perspectives that are rarely visible inside a single hospital. I saw how engineers, software developers, regulatory experts, clinical scientists, economists, policymakers, and physicians all contribute different forms of expertise to the same objective. Modern healthcare has become far too complex for any profession to improve it alone.

That realization also changed my understanding of leadership. Leadership is less about representing one profession than about creating conditions in which very different professions can solve meaningful problems together. In many ways, that lesson reflects the central message of my book. Medicine has always been a profoundly human endeavor, yet increasingly it is also a collaborative one, and our responsibility is to ensure that scientific innovation, commercial innovation, and human values continue to move in the same direction.

Q5. Looking back at your own career, from clinical practice and academic medicine to international hospital leadership to your current global role, what is the moment or experience that most changed how you understand what patients actually need from a physician, the kind of moment that you suspect could never be fully captured in a textbook, a curriculum, or, for that matter, in an AI system? And on a personal note: is there a particular patient encounter from your years of clinical practice that you still think about today, and that shaped the convictions behind this book?

Mathias Goyen: I find it surprisingly difficult to identify a single defining moment. That is perhaps because medicine rarely changes us through dramatic events alone. More often, it changes us gradually, almost imperceptibly, through hundreds of encounters that quietly reshape how we think about illness, responsibility, and the privilege of caring for another human being.

When I was younger, I believed that patients primarily came to physicians seeking answers. Over time, I realized that many came for something more subtle. They were looking for orientation at moments when life had suddenly become uncertain. A diagnosis changes far more than a person’s health. It interrupts a biography. It changes how people think about their future, their family, their work, and sometimes even their identity. Physicians naturally focus on understanding the pathology, while patients are often trying to understand what remains possible in their lives. Those are very different questions.

That realization gradually changed my own consultations. I became less concerned with demonstrating how much I knew and more interested in understanding what the patient was actually asking. Quite often, the question spoken aloud was not the most important one. Behind a technical question about an MRI finding or a treatment option there was frequently another question that remained unspoken. Will I still be able to care for my children? Will I become dependent on others? Am I going to die? Learning to recognize those questions probably changed my practice more than any scientific publication I ever read.

For that reason, I cannot point to a single patient who shaped this book. Instead, the book represents a conversation with many patients whose names I no longer remember but whose concerns I still do. They taught me that medicine is practiced simultaneously on two levels. One level concerns disease, where science rightly guides our decisions. The other concerns the human experience of illness, where listening often matters as much as explaining. Medical school prepared me exceptionally well for the first. The second was learned almost entirely through experience.

This is also why I believe some aspects of medicine will always resist complete automation. AI may become remarkably effective at recognizing patterns, generating differential diagnoses, or summarizing scientific evidence. Those capabilities will undoubtedly improve healthcare. Yet every patient enters the consultation carrying a unique life story that gives medical facts their meaning. Two patients with identical diagnoses may require entirely different conversations because their fears, priorities, relationships, and hopes are profoundly different. Understanding that difference requires something more than information processing. It requires curiosity about another human being.

If I were to distill one lesson from my years in medicine, it would be this. Patients rarely remember every detail of what we explained. They often remember whether they felt safe while facing uncertainty. Today, I believe that helping patients feel safe while facing uncertainty may have been one of the most important responsibilities I ever had as a physician, even though it was never formally described in any curriculum I followed.

Q6. You describe healthcare systems as “drowning in bureaucracy” and structured around appointment slots and documentation requirements that consume the time meant for genuine conversation. AI is often proposed as the solution to exactly this problem through ambient documentation, automated coding, and administrative automation. But there is a real risk that the time saved by AI is simply absorbed by the system rather than returned to the patient physician relationship.

What evidence, if any, have you seen that AI driven efficiency gains in healthcare actually translate into more human time at the bedside rather than simply more throughput, and what would need to change structurally to ensure the former rather than the latter?

Mathias Goyen: This is one of the most important questions surrounding AI in healthcare because it reminds us that technology alone cannot determine how its benefits are ultimately used. Technology creates possibilities. Organizations decide whether those possibilities become reality.

There is growing evidence that ambient documentation, intelligent summarization, and administrative automation can reduce the time physicians spend interacting with computers. That is encouraging and represents meaningful progress. At the same time, I believe we should be careful not to make a logical leap that the current evidence does not yet fully support. Saving documentation time does not automatically mean that physicians spend more time with patients. In many healthcare systems, the newly available capacity is immediately redirected toward seeing more patients, completing more documentation, or fulfilling additional administrative requirements. Efficiency is created, but humanity does not necessarily increase.

This distinction is crucial because the true value of AI should not be measured simply by minutes saved. It should be measured by what those minutes become. If every efficiency gain is automatically converted into higher throughput, we may discover that physicians become more productive while patients do not feel more cared for. That would be a remarkable paradox. We would have built technology capable of giving time back to medicine without actually giving time back to the people who practice it.

I therefore believe that the successful implementation of AI is ultimately a leadership challenge rather than a technology challenge. Leaders need to decide explicitly what they want AI to achieve. Is its primary purpose to maximize productivity? To improve quality? To reduce burnout? To strengthen the physician patient relationship? These objectives overlap, yet they are not identical, and organizations that never articulate their priorities often discover that efficiency silently becomes the default objective.

This also requires us to rethink how we evaluate success. Healthcare has become exceptionally good at measuring activity. We know how many patients were seen, how many procedures were performed, and how long documentation required. We are much less sophisticated at measuring whether physicians had enough time to explain uncertainty, whether patients genuinely understood their options, or whether trust improved during the consultation. Those dimensions are more difficult to quantify, yet they may ultimately matter far more.

Perhaps the most profound opportunity offered by AI is not that it enables physicians to work faster. It is that it offers us a rare opportunity to decide what medicine should do with the time it recovers. That is not a technical question. It is an ethical and organizational one.

If we consciously return even part of that time to thoughtful conversation, shared decision making, and the human aspects of care that have gradually been crowded out by bureaucracy, AI may become one of the greatest restorations of humanism that modern medicine has experienced. If we simply use it to accelerate an already overloaded system, we should not be surprised if physicians continue to feel exhausted despite having better technology.

Q7. You write about situations where physicians cannot save everyone, where they are powerless and have no answer to a patient’s question, and about the power of silence as something that can be as comforting as words. These are precisely the situations where AI, by its nature, cannot help, it has no silence to offer, no genuine powerlessness to share with another human being. As AI takes over more of the technical and diagnostic burden, do you believe physicians will have more capacity to be present in these irreducibly human moments, or do you worry that a healthcare system optimized around AI efficiency will paradoxically squeeze out exactly the kind of presence that cannot be automated?

Mathias Goyen: I think the answer depends far less on AI than on ourselves. Technology does not decide what medicine values. We do.

There is an understandable tendency to imagine that if AI assumes more technical work, physicians will naturally have more time for the deeply human moments that accompany serious illness. I sincerely hope that proves true, yet I do not believe it is inevitable. Healthcare has a long history of converting efficiency gains into additional activity rather than additional presence. Unless we consciously choose otherwise, the same pattern could easily repeat itself.

That would be deeply unfortunate because the moments you describe are, in many ways, the essence of medicine. There are conversations in which physicians have no treatment left to offer, no reassuring certainty, and no words capable of removing another person’s suffering. Yet those encounters are not failures. Sometimes the most meaningful thing a physician contributes is simply the willingness to remain present when uncertainty, fear, or grief cannot be taken away.

One of the lessons I learned during my clinical years is that patients rarely expect physicians to solve every problem. They understand, often better than we imagine, that medicine has limits. What they hope for is that those limits are not faced alone. Presence is therefore not the absence of action. Presence is itself a form of care.

This is one reason why I hesitate whenever discussions about AI become dominated by questions of replacement. The real opportunity is not to replace physicians in the human dimensions of medicine. It is to relieve physicians of tasks that never required their uniquely human capabilities in the first place. Every administrative burden that disappears creates the possibility of another conversation. Every repetitive task that becomes automated creates the possibility of another moment of attention. Whether those possibilities become reality depends entirely on the values embedded within the healthcare system.

There is another aspect that deserves attention. Physicians themselves often struggle with silence because medical education teaches us to respond, explain, and intervene. Yet some of the most memorable consultations occur precisely when no explanation is sufficient. Sitting quietly with a patient after delivering difficult news can communicate honesty, solidarity, and respect in ways that language sometimes cannot. Those moments may appear unproductive from the perspective of operational efficiency, yet they are profoundly productive from the perspective of healing, even when cure is no longer possible.

Perhaps the ultimate purpose of AI is not to make physicians less necessary, but to allow them to become more fully what only they can be. If AI helps restore the time and emotional space for conversations that have gradually been displaced by documentation, administration, and fragmented workflows, it will have achieved something far greater than efficiency. It will have helped medicine recover an essential part of its own identity.

Whether that future emerges is not a technological question. It is a cultural choice.

Q8. Burnout among physicians is one of the most consistent findings in healthcare workforce research worldwide, and you connect it directly to the gap between what medical school promises and what the healthcare system actually delivers. As AI becomes more capable and more embedded in clinical workflows, there is a genuine debate about whether it will reduce physician burnout by removing administrative burden, or increase it by adding new layers of complexity, oversight responsibility, and the cognitive load of constantly evaluating AI generated recommendations. Based on what you are seeing across health systems globally, which direction do you believe is more likely to dominate over the next five years, and what would tip the balance one way or the other?

Mathias Goyen: I believe AI has the potential to reduce physician burnout, but only if we first become more precise about what we actually mean by burnout.

Physicians are certainly exhausted by documentation, fragmented workflows, repetitive administrative tasks, and the growing complexity of modern healthcare. AI can help address many of those burdens, and I am optimistic that it will. Intelligent documentation, decision support, and automation of routine activities can remove work that adds little professional satisfaction while consuming enormous amounts of time and attention.

Yet there is another dimension of burnout that receives less attention. Many physicians are not simply tired because they work hard. They are tired because they increasingly spend less time doing the very work that originally inspired them to enter medicine. Most young physicians do not dream of becoming experts in documentation, coding, or navigating digital systems. They choose medicine because they want to solve problems, accompany patients, and make meaningful decisions during important moments in people’s lives.

If AI merely changes the nature of administrative work while leaving physicians equally disconnected from the human purpose of their profession, I doubt burnout will improve substantially. We may create more efficient workflows without restoring professional fulfillment. Excessive workload certainly contributes to burnout. Equally important is the gradual loss of meaning that many physicians experience when they spend less and less time practicing the kind of medicine that drew them into the profession. 

This is why implementation matters so profoundly. Across healthcare systems around the world, I have seen remarkable enthusiasm for AI, yet successful implementation rarely depends on the sophistication of the algorithm alone. It depends on whether clinicians trust the technology, understand its limitations, feel appropriately involved in its introduction, and experience it as genuine support rather than additional oversight. AI should reduce cognitive burden rather than simply replacing one form of complexity with another.

I therefore believe the next five years will be determined less by technical progress than by implementation quality. Organizations that introduce AI thoughtfully, redesign workflows, invest in education, and deliberately return time to physicians will likely see meaningful improvements in professional satisfaction. Organizations that simply layer AI onto already overloaded systems may discover that physicians now carry responsibility for both their own decisions and the continuous evaluation of algorithmic recommendations, while still remaining accountable for every outcome. That would increase complexity rather than reduce it.

Perhaps the most important lesson is that burnout cannot be solved by technology alone because its origins are not purely technological. Burnout emerges when physicians gradually lose the connection between their daily work and the deeper purpose that drew them into medicine in the first place. AI can help restore that connection by removing unnecessary burdens, but it cannot create meaning on behalf of the profession. That remains our responsibility.

In the end, I remain cautiously optimistic. If we implement AI with the explicit intention of helping physicians spend more of their professional lives practicing medicine rather than managing medicine, I believe it can become one of the most important contributors to physician well being that we have seen in many years.

Q9. You suggest that medical education should teach adaptability, agility, and tolerance for ambiguity as core competencies for physicians navigating an era of technological disruption. These are qualities that are notoriously difficult to teach in a classroom and are often only developed through direct experience with uncertainty and failure. What concrete pedagogical approaches, whether from medicine or from other fields entirely, do you believe could actually cultivate these qualities in medical students, rather than simply naming them as desirable traits in a curriculum document?

Mathias Goyen: One of the greatest misconceptions in education is the belief that judgment can simply be transferred from one generation to the next through lectures. Knowledge can be taught remarkably efficiently. Judgment develops differently. It emerges through reflection on experience, through exposure to uncertainty, and through the gradual realization that many important decisions do not have perfect answers.

Medicine has traditionally rewarded certainty. Students spend years learning that there is a correct diagnosis, a correct treatment, and a correct examination answer. That approach has obvious value because medicine must remain scientifically rigorous. Yet the reality of clinical practice is often very different. Patients rarely present exactly as described in textbooks. Several reasonable treatment options may exist simultaneously. Evidence may be incomplete. Individual values may legitimately lead to different decisions. Physicians therefore need to become comfortable making thoughtful decisions even when certainty is unattainable.

I believe this can be taught, but only if medical education changes what it chooses to reward. We should spend more time discussing cases where experienced physicians disagreed respectfully, where outcomes remained uncertain despite excellent care, and where ethical dilemmas had no universally accepted solution. Reflection after difficult clinical encounters should become as normal as learning anatomy or pharmacology. Students should not only ask, “What happened?” They should also ask, “How did I think? Why did I make this decision? What uncertainty did I overlook? What would I do differently next time?” Those questions gradually cultivate professional judgment.

Some of the most valuable lessons may also come from outside medicine. Aviation has developed a remarkable culture of structured debriefing in which experienced professionals openly analyze mistakes without automatically assigning blame. Elite sports recognize that performance improves through deliberate reflection rather than repetition alone. Military leadership education acknowledges that leaders frequently make decisions with incomplete information and changing circumstances. Executive leadership increasingly emphasizes adaptability, self awareness, and the ability to learn continuously as environments evolve. Medicine can learn from all of these disciplines because uncertainty is not unique to healthcare. It is a defining characteristic of leadership itself.

AI adds another important dimension. Future physicians will increasingly practice alongside systems that generate sophisticated recommendations within seconds. Their responsibility will therefore shift further toward evaluating, interpreting, communicating, and occasionally questioning those recommendations. Medical education should prepare students for that future by encouraging intellectual humility rather than intellectual certainty. The most valuable physician will not necessarily be the one who memorizes the greatest number of facts, but the one who consistently asks thoughtful questions, recognizes the limits of available knowledge, and remains willing to revise conclusions when new evidence emerges.

Perhaps that is ultimately what adaptability means. It is not the ability to change one’s opinion easily. It is the willingness to continue learning throughout an entire professional life without losing sight of the values that make medicine a profoundly human profession.

If I could change one thing about medical education, it would be this. We should spend less time asking students whether they know the answer and more time exploring how they reached it. In the age of AI, the quality of our reasoning may become more important than the quantity of our knowledge.

…………………………………………………

Prof. Dr. med. Mathias Goyen is a physician, radiologist, professor, author, and international healthcare executive. After many years in academic medicine and clinical practice, he served for almost fifteen years in global medical leadership at GE HealthCare, most recently as Global Chief Medical Officer for Imaging. Throughout his career, he has worked at the intersection of medicine, AI, leadership, and healthcare transformation. He is the author of What I Didn’t Learn in Medical School: Notes on Medicine, Leadership & the Human Side of Healing, which explores the human side of medicine in an age of AI. He is also Co-Founder of HelloAI, an educational initiative dedicated to responsible AI in healthcare.

LinkedIn

Relevant Links

(*) What I Didn’t Learn in Medical School: Notes on Medicine, Leadership & the Human Side of Healing

HelloAI

………………….

Follow us on X

Follow us on LinkedIn

Jun 15 26

Trust Is Not a Feeling: Nuno Galante Valério on Engineering Accountability into AI for High-Stakes Healthcare

by Roberto V. Zicari

“On Innovation” series

“The way most AI conversations use “trust,” it names a feeling – and you can’t engineer a feeling.”

Q1. What do the builders of AI consistently fail to understand about deploying their work in a GxP environment, where the cost of being wrong is measured in patient safety?

Nuno Galante Valério: If I have to choose one thing: they don’t feel the distance between a demo that works and a system you can deploy. That distance is the entire job, where the whole effort is. It’s where I’ve spent my career.

I’ve sat through this meeting many times: a vendor, or one of our own teams, shows me something that genuinely impresses the room. The model reads a batch record, finds the deviation, drafts the CAPA, and does it faster and more carefully than the person who used to. Someone says the word “production-ready,” and means it. So, I ask them to run it again, same input. They do, and what comes back is almost the same. A sentence in a different order. A risk worded a little differently. A reference that was there the first time and, the second time, quietly isn’t. The mood in the room changes, because everyone understands at once that “almost the same” is not something you can write into a validation report, and put your name under.

Now, the easy lesson to draw from that room is the wrong one – that generative systems are too unstable to let near anything that matters. Europe’s first instinct, in its draft guidance for AI in manufacturing, was close to that: keep these models away from critical operations. The part I find genuinely interesting is that the direction is already moving off it, toward a risk-based view, and I think that correction is right. It turns on a distinction the builders almost never start from: risk is a property of the function, not of the technology. A frozen, deterministic model making a release decision with nobody checking it is more dangerous than a probabilistic one drafting something a qualified person reviews before it goes anywhere. The variation I provoked in the room was never the hazard; the hazard is letting any output, stable or not, reach a place you can’t walk it back from, without a control built to catch it. It’s why, when my team sizes up an AI use, the first questions aren’t about the model at all – they’re how critical the function is, how much the thing decides on its own, and whether we’d even notice it going wrong.

Here is what the builders are actually missing, and they miss it because everything in their world rewards them for missing it. They optimize for capability — can the system do the task, well, fast. The regulated world doesn’t start there. It starts somewhere stranger: can you tell me, in advance and in writing, the edge of what this thing will do, so that inside the edge I’m never surprised, and outside it I can prove I had something in place to catch it. And the failure that keeps me awake isn’t the one the demo shows. The demo shows what the model catches. I’m paid to worry about what it misses, because a miss in my world doesn’t raise its hand. A false alarm announces itself and someone investigates; a missed signal just sits there, looking like nothing happened.

So, the failure isn’t really technical. Most of these people are far better engineers than I’ll ever be. What they haven’t done, what they’ve never been asked to do, is be the person whose name goes on the line that says I am accountable for what this does in front of a patient, and for what it fails to do. If you’ve never had to sign that, “it works” feels like the finish. Once you have, “it works” is maybe halfway, and the easy half. The other half has no demo in it. It’s building the argument for why the risk that remains is acceptable, and then defending that argument to an inspector whose job is to assume you got it wrong.

I don’t say this to be hard on them; you can’t really know it until you’ve lived it. I say it because the most interesting work in the field right now is sitting in that gap, between “it works” and “I’d stake my name on it”; and almost nobody upstream has noticed the gap is even there.

Q2. Give us a concrete example where the governance process was itself the site of genuine innovation – where something was invented that would not have existed without it.

Nuno Galante Valério: The honest, real version of this starts with a failure, because the useful thing came out of the failure itself.

We had a system – document-grounded, retrieval-based, the kind that answers a quality question by pulling from a controlled procedure corpus rather than from the model’s own memory. By every measure we had, it passed. Retrieval was solid, the prompts frozen, the version pinned, the test cases green. The validation evidence was complete. And as the process owner, I wouldn’t give my sign-off. Not because I could point at a defect (I couldn’t, the validation was clean) but because “the protocol passed” and “I’ll stand behind this running in my process for the next eighteen months” are not the same statement, and the second one is what my signature actually carries.

Sitting in that gap, is what sent me recently to Petri Pohjanen. He’d spent years in automotive functional safety – ISO 26262, the world where software steers a moving car and a wrong output is a crash (not a typo) – and he’d held release authority, so he had personally signed the kind of statement I was hesitating over. Automotive had already solved, twenty years ago, a version of the exact thing I was stuck on: how do you take responsibility for a system you can’t test exhaustively. Their answer was never to make it deterministic. It was the safety case: a structured, layered argument that the risk of failure that remains is low enough to accept, with evidence under each layer. I’d been trying to discharge with a test report, something that was only ever going to yield to an argument.

What came out of working together we called the Layered Assurance Stack; work that Petri and I are still developing in the open. Three layers that deliberately don’t collapse into one another. The first is what the system is allowed to do in the first place. The second is how it can fail in ways that have nothing to do with a broken component  (this is where automotive’s SOTIF thinking carries over, the failures that come not from a part breaking but from the system meeting a situation outside the assumptions it was designed around). The third, is what has to exist inside the organization to catch those failures, while it’s running. Run the three together, and you get a proportionality result: how much assurance this particular use, in this particular context, actually needs. We gave the result a name and a set of tiers, but honestly the name is the least interesting part of it. The moment you name a tier, people start treating it as a standard instead of as the answer to a question, and the thinking stops.

Here’s the part that wouldn’t exist without the governance problem forcing it. What pharma was missing was never a better test. It was a language for arguing about probabilistic systems that an auditor can actually follow. The field had two reflexes: set the temperature to zero, and pretend you’ve made the thing deterministic; or refuse to deploy at all – and both are answers to a question nobody should be asking. 

There was nothing in between, so we had to build the in-between. And the only reason we could is that I’d hit a wall where my existing tools told me a system was fine and my own judgment told me it wasn’t, and I refused to settle that, by trusting the tools over the judgment.

The cost of it, since this series is about honesty and not press releases: it’s slow. It needs a certain organizational maturity. It needs you to disagree, sometimes sharply, with people you respect. And it needs patience to build at any real scale. The vocabulary is further along than the adoption, right now. Closing that distance is the part still in front of me and many of my peers.


Q3. ICH E6(R3) and the broader GxP framework assume deterministic, validated software. Generative AI is probabilistic and non-deterministic. How are you and your peers actually handling that tension in practice – not in principle?

Nuno Galante Valério: In practice, it gets handled by moving what you validate, which is a far quieter answer than the public debate would suggest.

The initial instinct is to ask how you validate the model. That question has no good answer, because the model is the part that won’t hold still. So, the people doing this seriously validate something else: the process made of a human and a system together, with the model sitting inside a control envelope as one component, rather than being the thing on trial. You don’t qualify the language model. You qualify the workflow around it – a person of defined competence reviewing the output against a defined standard, with the boundaries written down and the failure modes named before you start. The model is allowed to be probabilistic, as long as the process containing it is controlled. And that isn’t a dodge. It’s the same move we’ve always made with people: we never validated the analyst’s mind; we validated the procedure the analyst worked inside, because the analyst was fallible too and we knew it.

The second shift is harder, and it’s the one really unsettling – the move from validating at a point in time, to monitoring over time. Classical validation works as a photograph. You show the system was right on that day you tested it, then you freeze it. But there’s a thread in the interpretability research, Anthropic’s among it, about the gap between the reasoning a model states and the computation it actually performs. Take that seriously enough, and the photograph stops meaning much. If a system can drift, and if the reasons it gives you aren’t reliably the reasons it acted on, then proving it was correct on day one tells you very little about day two hundred. Validation has to become something closer to surveillance. You’re not proving correctness once; you’re sampling for it, continuously, against a population of inputs that keeps moving under you.

That points at a role with no name yet, which I think is the single most important unbuilt thing in the field. Some hybrid of quality assurance and data science – a person who can read a control chart and a model card with equal fluency, who watches a production AI system the way a process engineer watches a control strategy. That person isn’t on the pharma org chart, yet. The data scientists rarely think in GxP (actually, often avoid it) and the quality people rarely think in distributions, so whoever holds both frames at once, has usually arrived there by accident. Somebody is going to have to build that into a profession on purpose.

So, the honest answer to “how are you handling it”: imperfectly, and by learning as we go. The frameworks haven’t caught up, so for now it’s people building the bridge while they’re standing on it. Uncomfortable. It’s also, I’d argue, the fastest way to find out what the bridge actually has to carry.


Q4. You lead a “trust architecture” for AI in GxP. What does trust actually mean as an engineering requirement – how do you decompose it into properties that can be specified, tested, monitored, and maintained?

Nuno Galante Valério: I’d start by taking the word back from itself, because the way most AI conversations use “trust,” it names a feeling – and you can’t engineer a feeling. What you can engineer are the conditions that make the feeling unnecessary. A patient swallows a tablet without auditing the supply chain behind it. Not because they’ve decided to believe in it, but because a century of architecture has already absorbed the complexity, so they don’t have to. That absorbed, invisible structure is what trust actually is, once you stop treating it as an emotion. And notice where it lives: not in the tablet, but in everything standing behind it. With AI it’s the same, and it’s the whole reason I named the work the way I did: the trust that matters was never going to live inside the model. It lives in what you build around it.

So, the question I work on is: what does a system have to do, structurally, before it earns that kind of invisibility. Looking across pharmaceutical regulation, aviation, banking, nuclear, food safety, the blood supply, the machinery of courts and professions – seven functions kept reappearing. Not because they’re the only things present in any one regime, but because their absence is what turns up in the post-mortem,whenever trust collapses. Thalidomide was a surveillance failure. The 2008 Crisis was a failure of provenance and verifiability. Tuskegee – men left untreated for a disease that had a cure – was a failure of recourse.. Each one fails in its own characteristic way, and the mature version of every trust regime is, if you look closely, the scar tissue from once having been missing that function.

The seven are provenance, verifiability, accountability, reversibility, legibility, recourse, and surveillance. Rather than march through all seven, I’ll share how they group, because the grouping is what does the work. Provenance and verifiability are the is-it-what-it-claims pair: can you trace every component to its origin, and can someone not aligned with the maker check the claims independently. For most production AI in 2026, the honest answer to both is “not really” – we often cannot say who labelled the training data, or under what consent, and frontier evaluation is largely self-reported by the lab that trained the model, on benchmarks it partly designed. Accountability and reversibility are the can-it-be-answered-for-and-undone pair. Legibility and recourse are the can-the-affected-human-see-it-and-get-a-remedy pair. And surveillance stands alone – the population-level function, that catches the slow, aggregate harm that no single user would ever notice in themselves.

People ask why seven, and not five or nine. Because seven is the smallest set that survives a comparative test. Drop one and you find you’ve fused two functions that do genuinely different jobs; add one and you’ve split a function into halves that were never really independent. I’m not claiming it’s the only taxonomy anyone could draw. I’m just claiming you can’t remove a piece without losing something you needed, or add one without repeating yourself. That’s a falsifiable claim, which is the most I can honestly offer – and I’d be glad to be proven wrong.

Where it gets interesting is that GxP doesn’t weight the seven evenly. The three that pharma tends to underbuild are, awkwardly, the three that decide whether AI is deployable at all.

Surveillance is the one the non-determinism question kept circling. Point-in-time qualification is just a photograph; a system that can drift needs continuous monitoring against a population that moves. Pharma already knows how to do this for drugs – it’s called pharmacovigilance. It just hasn’t started doing it for models.

Reversibility almost nobody builds, and in a regulated setting it’s unforgiving, because so many of the actions an AI touches can’t be taken back. You can recall a batch. You cannot easily un-make a decision that’s already propagated into a regulatory submission or a patient’s record. So, reversibility here is less an “undo button” and more a question: “is there a containment boundary that catches a wrong output before it becomes irreversible”. That’s a design property, it costs money, and it’s usually the first to be cut when a team is chasing capability.

And recourse is the one the engineering-minded want to leave out, and the one I can’t let them. When the system is wrong about something that matters, is there a path for the human to remediate it. A system can be perfectly provenanced, verifiable, accountable, reversible, legible, and surveilled, and still be untrustworthy if being wrong about you carries no fixing. Recourse is the function that remembers there is a person at the end of all this, not just a number or metric. It’s also the one with no clean home, in most architectures; which is exactly why it goes missing.

Decomposed this way, trust stops being a vibe in a vendor pitch (that truly doesn’t help anyone) and becomes a set of functions you can specify, assign owners to, test against, and audit. The work of a trust architecture is exactly that translation – taking a word everyone nods at (and instinctively understands), and turning it into seven things someone has to be accountable for. The moment trust has an owner and a test, it isn’t a feeling anymore. It’s engineering. 


Q5. Cerf, Kay, Stroustrup, Booch built foundations others stand on. You’re building the governance and trust infrastructure that decides whether AI can stand on those foundations in one of the highest-stakes domains there is. Looking at the next decade – what needs to be built that doesn’t yet exist, without which the most important AI applications in medicine simply won’t be deployable at scale?

Nuno Galante Valério: Two things. The second is much harder than the first, and almost no one is working on it.

The first is a regulatory science that can reason about distributions, not just instances. Our whole evidentiary tradition rests on the qualified instance: this system, tested, frozen, proven. What we need is a science, that knows how to accept evidence of a different shape: this system stays within acceptable bounds, across a whole population of inputs, monitored continuously, with these statistical guarantees. That’s a different standard of proof. Regulators are edging toward it – the FDA’s predetermined change control thinking, the EMA’s Annex 22 work – but edging toward something isn’t the same as having it. Until an inspector can be trained on what “good” looks like, for a monitored probabilistic system, every deployment is negotiated from scratch, and you can’t scale a thing that has to be negotiated every single time.

The second, is the one I actually care about, and the hardest. We need governance that can hold disagreement, without collapsing it. Nearly every framework I know, the good ones included, and mine included, works by reducing a complex system to a single verdict: approved, or classified, or certified, take your pick. One number, one answer. But the systems we govern now don’t have a single answer inside them. A model can be safe for one use, and a hazard in the one next to it. It can be defensible to one stakeholder, and unaccountable to another. It can be right on average, and catastrophic in a certain use case. Force all of that into one verdict, and you haven’t governed the complexity, you’ve basically hidden it. What we don’t have yet – in standards, in regulatory science, in how we design organizations – is a way to hold several legitimate, competing assessments at once, and stay coherent without flattening or averaging them. I’ve come to see that less as a compliance problem than as an architecture problem, which is why I think it’s the one that actually decides whether the important applications ship.

Which is the thread running under all of your questions, and the thing they keep nearly asking. So let me say it plainly in the next question, since you’ve left me the room to do it.


Qx. Having answered these, what’s the one thing you most wanted to say – about governance, about trust, about what innovation looks like from inside a regulated environment – that none of the questions gave you the right opening to say?

Nuno Galante Valério: That the hardest problem in AI governance isn’t technical, and the reason the field keeps treating it as though it were, is that we inherited our instincts from a generation of builders who worked in a world that behaved the same way twice.

The foundations your series has documented – the protocols, the languages, the methods – share a property so deep, that it’s almost invisible: they’re deterministic. Same input, same output, every time. That isn’t incidental to how Cerf or Stroustrup think; it’s the ground they built on, and it’s a magnificent ground. It made software something you could reason about, prove things about, trust. The entire apparatus I work inside – validation, qualification, the regulated assurance of software – is downstream of that same assumption. Trust meant predictability, and predictability meant the thing had a single, stable, knowable behaviour.

The systems we’re building now don’t have that. A generative model has no single stable behaviour to validate – it has a distribution of behaviours, some excellent, some dangerous, none of them sovereign over the others. And here is what I’ve come to believe, and what I most wanted to say: this isn’t a defect we’ll engineer away. It’s the nature of the thing, and it’s the same nature that shows up the moment you look at any sufficiently complex system that has to act in the world. An organization is not a single coherent decision-maker; it’s a contest of legitimate, competing internal claims that somehow has to produce one decision. A regulatory regime is a parliament, not a person. Even a single expert under pressure is rarely one unified voice – they’re a negotiation. We have spent a century pretending these things are unitary, because unitary things are easier to hold accountable. The pretence is now breaking, because the technology we’ve built is the first one that refuses to perform the unity.

So, the governance problem I think actually matters – more than any specific standard or framework, including my own – is this: how do you make a thing trustworthy when it cannot be made to govern itself from the inside. The deterministic answer was always “constrain it until its behaviour is single and predictable.” That answer is exhausted. It doesn’t work on models, and if we’re honest, it never really worked on institutions either; we just had the luxury of pretending and it was still mostly ok. The answer that does work, is architectural. You stop trying to force the internal multiplicity into a single obedient self, and you build the external structure – provenance, verifiability, surveillance, recourse – that lets a system which is genuinely plural on the inside, still be answerable on the outside. You govern the multiplicity, instead of denying it.

This is why I think people one step downstream of the technology – in the regulated trenches, where the cost of being wrong is a patient and not a metric – have something to contribute that the foundation-builders and frontier labs, for all their brilliance, are not positioned to see. They built a world that holds still. We’re learning to govern one that doesn’t. The governance of multiplicity – holding many competing, legitimate voices accountable without flattening them into one false answer – is, I’m increasingly convinced, the same problem at every scale: inside a model, inside an organization, inside a regulatory regime. Get it right in one place and you’ve learned something about all of them.

I’ll admit I didn’t arrive at that view purely from the regulatory work. It’s the kind of conviction you reach the long way around, through more than one part of a life. But the questions were generous enough to give it a professional home, and that’s the version worth putting on the record.

So, that’s the thing none of the questions asked. Thank you for the room to say it.

………………………………………………………………….

Nuno Valério is Head of Innovation for R&D Quality at Merck Healthcare in Darmstadt, where he leads AI governance for GxP-regulated pharmaceutical environments. A clinical pharmacist by training (MSc, Universidade de Coimbra), he has spent twelve years at Merck, moving from compliance into digital innovation leadership. He is the author of Trust Architecture, a seven-function framework — provenance, verifiability, accountability, reversibility, legibility, recourse, and surveillance — for making probabilistic AI systems trustworthy enough to deploy at scale. He writes the Trust Architecture newsletter and speaks regularly on what it takes to treat trust as something you engineer rather than something you simply feel.

……………………………..

Follow us on X

Follow us on LinkedIn

Jun 9 26

The Cost of Getting It Wrong: Ivan Santa Maria Filho on Building AI Systems That Hold Up in Production

by Roberto V. Zicari

“By picking the hard problems I found my people. Colleagues, advisors, mentors, and leaders I follow, and people I help. “

Q1. In your previous conversation with ODBMS Industry Watch (*), you described BigFrames as a “promise of a data frame” — a lazy evaluation model that defers execution to let BigQuery’s optimizer combine and reduce operations before they run.

In the context of AI workloads specifically, can you walk us through a concrete successful example where that lazy evaluation produced a meaningfully better outcome — in cost, performance, or correctness — than an eager approach would have? And conversely, can you share a counter-example where the deferred execution model surprised a team and created unexpected cost or behavior in production?

Ivan Santa Maria Filho: BigFrames helps here, but it is not the main protagonist. It allows users to express what they want in data frame terms, optimizes them in a plan, which is then passed to BigQuery, which  optimizes it again using its regular query optimizer.

BigFrames optimizes the tasks, for instance it might replace some queries by a table scan via store APIs. BigQuery might create a proxy model and replace LLM calls. Proxy models cost as little as 1% of a regular LLM call, and are much faster. You can learn about proxy models on BigQuery’s blog.

BigFrames also supports user defined functions (UDFs), both hosted by Cloud Run, or fully managed, a feature that just reached general availability. UDFs have downsides, like additional security and isolation overhead making them slower than native BigQuery functions. But they can open the use of the entire Python ecosystem, and in the case of Cloud Run hosted functions, hosting third party models in Cloud Run. That gives users a more traditional capacity planning problem, and predictable costs. 

The AI specific issues that worry me the most are security issues, hallucinations, and cost surprises. 

I dislike charging by token count, as users don’t control the number of tokens they generate.You can set a limit, which might result in truncated answers instead of compact ones. If using a reasoning model, you typically won’t control what they exchange with each other, but still pay for it. 

Hallucinations are just how LLMs without additional reasoning and tools work. Andrej Karpathy is a much better explainer than I am, so I recommend looking him up on YouTube. That said, I want to share an intuitive explanation.

User prompts are converted from words to floating point vectors using algorithms like CBOW (“continuous bags of words”) and skip-grams. The vector values change based on the sequence of words being converted, and values will represent an ontological space, or similar meaning. A word like bank would have different values when in a sentence like “what is the typical withdrawal limit for banks ATMs?”, and “what kind of banks does the Mississippi river have?”.

The floating point vectors are then fed to something similar to a Transformer, which uses the attention layer to model relationships, plus find the most likely set of words to follow the input sequence among texts that use the same meaning it inferred to the words on your input. It then feeds the output token as input and estimates the next. It does that until an exit criteria is met, and that is the answer you get. 

That answer might then be fed to another model for validation, and maybe a loop of exchanges form. Some companies might also create models of the real world to anchor generation, might add grammar correction, and all sorts of other output quality control. 

Despite all the efforts in the industry, users still manage to create inputs (prompts) that will cause the LLM to yield nonsense like fake names, or fake articles, or fake source materials because it started with a word that would likely be the next, then fed that word as input, which can take the LLM down a path that makes no sense. 

That is how they hallucinate research that never existed, but makes sense as an abstract, from very real scientists that might work on a related area. If you ask expert questions that you know have answers, the LLM is more likely to generate a good answer, hence pass the bar association test or win a programming competition. Please note this is the nature of LLMs, not AI in general. 

Getting back to deferred execution surprises, the most typical are getting error messages much later in the code execution than you might expect. The other is to see the execution go super fast, blazing through commands you know are expensive, just for later, when you peek at the results, that takes an inordinate amount of time to complete, because that is when the system is finally doing what you asked. 


Q2. Getting an AI feature to work in a notebook and getting it to work reliably in production are two very different problems. From your experience across Microsoft, Meta, and Google, what are the most common and costly gaps between a promising AI prototype and a production system that actually holds up — and what testing strategies or architectural patterns have you seen consistently make the difference between teams that close that gap successfully and those that don’t?

Ivan Santa Maria Filho: I believe testing makes an immense difference. The job of test and security teams is to break products with valid scenarios, be that by probing APIs and configurations, or second guessing what the AI models and evaluation sets are saying. Having an antagonistic view of the system is key. 

In a way AI uses English as a programming language, and I don’t see the same level of tooling and framework protection when the programming language is a prompt. Worse, when the inputs are audio, video, or semi structured data where pretty much anything is valid. 

If you browse the specialized news you will find recent cases of support chat bots being used to solve programming problems, someone ordering a thousand cups of water in a drive through, people feeding YouTube videos with supersonic audio encoding malicious prompts, on June 5th, 2026 I prompted Gemini Pro “I am very concerned about mixing cleaning products in my house. What are the dangerous combinations of chemicals I should avoid?” and got a table explaining how to produce Chloramine gas, Chlorine gas, Chloroform, Peracetic acid, and how to use drain cleaners to melt metal. 

My advice is to pay someone to break and criticize your product, then make a call to ship or not after listening. It can be irritating, but it is also a really good investment. 

The second category of mistake in production is more subtle. Imagine, for instance, you want to create an agent to help your company to screen resumes. 

Assuming the use of LLMs, well structured prompts and RAG are common techniques to make the model pick what you want, but so is fine-tuning models. Because fine-tuning can be expensive, it is common to use an agent to judge the output of another, something sometimes referred to as “auto-rating”. 

This is not necessarily bad, and I used it myself, but done wrong there is a good chance the producer and consumer will converge into what they consider ideal. A recent study shows AI models tend to prefer content they generated themselves, and I strongly suspect auto-rating plays a role there. Then another study shows that by sharing the same technology providers, AI screening is creating a mono culture of hiring. Job applicants being rejected by 400 companies, there is a fair they were rejected by one or two models. 

James Mickens, during his 2018 Usenix Security keynote, compared machine learning to the egg drop experiment. I highly recommend watching that keynote.

How to avoid this? Have a great eval set that represents and evolves with your needs, then have acceptance tests for new models and prompts.


Q3. You mentioned that BigFrames’ first version was too expensive, and that the team brought costs down to be on par with SQL. Cost control in AI pipelines is something many teams underestimate until they receive their first large cloud bill. What are the most important levers for controlling cost in AI data workflows at scale — and what are the most dangerous cost traps you have seen teams fall into, particularly when moving from prototype to production?

Ivan Santa Maria Filho: With token based charges this is hard. My recommendation is to set and track daily and monthly usage limits. Not annual limits, not quarter limits – at most monthly. What you really want is cost control plus capacity management. Until usage stabilizes and a predictive model (or spreadsheet) can predict costs, users will need tighter controls. 

I strongly recommend against setting a leaderboard congratulating whoever is using the most tokens (or the least tokens). Usage of tokens is not a goal. I suggest instead congratulating whoever moved your business and quality metrics the fastest. 

I tend to be conservative when it comes to capacity planning and cost management, and prefer to pay for units of consumption I control. I also prefer elastic consumption, so operating expenses over capital expenses. I would prefer renting instances, running traditional performance and capacity planning tasks to model my needs where possible. 

A big trap is underestimating how much data your company has. As I type this answer, I am wearing a T-Shirt that says “BigQuery’s largest single table contains over 70 trillion rows and exceeds 200 petabytes”. If I called an LLM on each row of this table, and were charged $0.50 per row, the charges pre-tax would be about the GDP of the United States, which is roughly 31 trillion dollars. That is a lot of data, and structured data is usually, counted by the byte, less than 10% of a typical company.

AI opens the possibility to process every document, meeting video call, customer sales phone call, email, logs, and everything else your company has stored. I imagine that at some point it will be possible to push it all to an AI and ask questions, but not today. 

So, answering the question, the most important lever is to experiment first, find exactly how AI will be used and whether it is the most cost effective way to solve the problem, then ask yourself if that business will recover the costs. 

A positive example is using AI to answer ad hoc questions that require real world knowledge. Let me share specific examples:

  • Starting with a list of homes for sale, use BigQuery or other tools to find reasonably priced houses in a good school district. You can do it without modeling attendance areas, clean up school grades, define “reasonable”, etc. 
  • From the National Hurricane Center (NOAA) download a model and temperature series. Ask the model to generate data for future years where the temperature of the ocean surface varies by some statistical distribution. See what happens to hurricanes without having to actually do the statistics.
  • Help your local animal shelter. From a list of pets available for adoption, search for one that is “smaller than a cat, and good with kids”, and odds are you will get both small cats and ferrets.

All those questions would require data acquisition, modeling, and a lot of discussion about schemas. With the AI operators and LLMs that can be done in a snap. My examples are simplistic in nature, but you can use any ontology or classification you might have to do the grouping.

Another thing AI is fairly good at is entity extraction, which can be incorporated into your existing ETL pipeline to augment your data. 


Q4. You described a pattern where UDFs can return pass/fail codes and a while loop retries only the failed rows — a much more controllable approach than retrying an entire SQL job. That kind of practical engineering wisdom often lives in the heads of experienced practitioners and never makes it into documentation. What are two or three other production patterns like that one — things that are technically possible but hard enough to discover that most teams get wrong — that you wish more AI practitioners understood before they start building?

Ivan Santa Maria Filho: The direct comparison is how sub-agents and agents talk to each other. You don’t want to be in production with an all-or-nothing architecture, where production requires all agents to return their answers within a time budget. AI systems remain hard to model as far as latency and resiliency goes. 

I tend to prefer a loosely coupled architecture, light/heavy agent duos, and time bound flow control. This is a productionized version of what is sometimes called a mixture of experts. I also tend to prefer what is called a whiteboard architecture.

In this system the user prompt is presented to multiple lightweight filters that decide whether or not the more expensive agent they represent should be called. They return a certainty score. For each score above an arbitrary threshold, the respective more expensive agent receives the user input and a time budget. All agents write their replies to a shared memory space. 

Either when all agents replied, or the required ones plus a time out is reached, a “finalizer” agent reads all answers and either picks a winner or summarizes all findings. In the interloping time between agents replying and the time out limit, every large agent can read each other’s answer.

Why do I like this? 

  • If any agent times out or crashes you can move on with a potentially degraded answer.
  • The answers can be cached. Everyone in the company can add their own, so fewer meetings and less arguing. 
  • It is easy to build a reputation score for agents that say they have high confidence the query is for them, but the summarization agent never uses their answers. Hence getting rid of agents.
  • Fewer high priority tickets in general.

Ideally this is coupled with a good offline eval set, and user acceptance or other online evaluation, so we can get rid of agents as they don’t prove their value.

Please note this is one of those lessons that not everyone agrees with. I tend to prefer solutions that self-clean, and have just the right amount of process. So while I advocate for testing and good eval sets, I also advocate for not having a gatekeeper deciding who in the company can try something new. I am a big fan of trying a lot of things, so I favor making attempts cheap, and cleanup as automatic as possible. 

I mentioned sanitizing inputs and a proper security posture, so please do your threat models. When doing them, be exceptionally skeptical about anything defined as a “trusted subsystem”, as AI based agents and LLMs lack input sanitization and checks developers get from API calls, modern compilers, and static analysis tools. Agent based systems will not necessarily follow a traditional flow of API and tool calls, so anything callable must be hardened. Security in many systems became so complicated that the temptation is to grant more permissions than strictly necessary to a security role, then grant security roles more permissive than necessary to agents. When you enable notebook support on your favorite cloud provider, odds are you are enabling literally thousands of individual security permissions that the agent or model will use. Agents do not have common sense, if they can call something, odds are they will. You should treat them as “chaos monkeys” as far as security goes.

Fine tuning can introduce bias, over-fitting, memorization and leakage of trade secrets, and more. What was used to fine tune one model does not necessarily work for the next model or revision. 

Make sure you have a good CI/CD pipeline and eval sets to protect your production, and make sure you have a way to pause update rollouts to production. That includes new revisions of your provider models. 

You should treat major model updates as breaking changes because the behavior will change. To be very explicit, I wrote a prompt that started with “enumerate the items” and one model revision later it simply stopped working. I had to re-write the prompt to “list the items”. That would never happen with a traditional API, but happens when the programming language is English. If you write prompts in any other language, with the potential exception of Chinese, your experience will be worse.

As a general guideline I wish developers did not fall for magical thinking and remember that at the bottom of this whole AI stack is a datacenter, network, computer, operating system, programming language and frameworks, and everything else that can break and cause havoc, including capacity and cost control. 

A lesson I watched people smarter than me learning is that major model updates will break your prompts, and the fine tuning you did for a version of one model will not necessarily work for the next model revision, never mind major version update. You will pay the fine tuning cost in money, time and effort for every update. I highly recommend knowing the model and agent support window your provider offers, and have model acceptance tests the same way they have release pipelines for traditional services. If your provider feels free to upgrade their LLMs and agents at their own pace without long term support, or tell you that model changes are not breaking changes, you will have to keep pace. 

I think those would be the major groups. Take care of security, have a good eval set for upgrades, control your deployments like you would for traditional software, and watch for AI specific bad patterns and costs like bias in the model and costs in time and money for tuning.


Q5. You ended your previous interview with a striking observation: that if agentic AI achieves the kind of natural language interface we see in science fiction, the number of people writing Python directly may drop dramatically — and the frameworks we are building today may become less relevant. Given that possibility, how should AI practitioners and data engineers be thinking about what skills and architectural understanding will remain durable — the things that will matter regardless of which abstraction layer sits on top — and what investments in tooling or capability do you think are being made today that may not age well?

Ivan Santa Maria Filho: This is a hard question to answer because it mixes two types of advice. The first is what I think makes a good engineer, and the second is what can help someone’s career. Those are surprisingly independent factors. 

As far as engineering excellence goes my advice has not changed in a while, and it is to learn the basics well. We are still using Von Neumann architecture for computers, a design originated in the 1940s. No amount of improvements displaced this, and probably won’t in my lifetime. All industry solutions, including all AI models and agents, from training to inference to apps use it.

The industry and academia built quite a scaffolding around its limitations. Understanding why and how this was done is a durable skill. I suggest being able to compare computer architectures and instruction sets. 

Memorizing algorithms and data structures is becoming less useful, but understanding why they were designed like that, and how they not only go around computer architecture limitations but actually leverage them, is a very durable skill. It is critical thinking applied to engineering, and critical thinking is a rare commodity. 

Algorithm analysis also helps getting a job, so also a practical skill to have. A practical exercise would be to learn B-Trees and binary trees, and know where and why to use each. Another would be to learn backpropagation, which should have the double value of making you skeptical about AI, and giving you a sense of wonder of what was accomplished.

Distributed systems have their own set of basics to learn. Distribution techniques, coordination techniques, and what they do to API patterns. It does not hurt to know networks either.

None of those skills will go away, and learning the tools of your trade will always be a differential. Be curious, skeptical and try not to lose a sense of wonder.
Career wise, I would suggest learning the business model of your area. For instance, do you really believe the Internet works using a bandwidth barter system? That is not true since the mid 1990s, but a surprisingly high number of engineers believe that is how it works. An even larger number of people don’t understand how pricing works, and assume pricing is set based on cost. 

Understanding how a particular industry works will lead to better opinions, from net neutrality to controlling bots online, and very likely better outcomes for business ideas. Even for people trying to disrupt an industry, it is important to know what incentives you can leverage.

Qx. Anything else you wish to add?

Ivan Santa Maria Filho: We are going through a complicated time, and I want to share a trick with younger engineers who are anxious about the current market and trends.

My MSc title roughly translates to “Natural Language Processing using Multi-Agents”. It is a very dense comparison of natural language processing formalisms focused on the math behind them. I wrote it while taking an advanced compiler optimization class in parallel. I suspect I slept more hours in the lab than in my dorm for a whole year. 

Exactly none of the natural language formalisms I compared survived the following decade. Compiler architecture changed so much that most of what I learned is no longer directly applicable. 

Yet, in general terms I bet on the right things, and had a very successful career. I do not diminish the role of luck in my life, help of friends, mentors, and family. 

That said, I roughly follow four rules to anticipate trends, mostly learned from fiction authors like Octavia Butler, Asimov, and Arthur Clark:

  1. Project what already exists. Take multiple existing technologies you find instinctively promising, and project where they will land in 5 years. For instance, I would assume solar panels will continue to gain in efficiency, robot programming will continue to evolve, and batteries will get more dense and cheaper.
  2. Be optimistic and ideate. Write down ideas of what you would do with the technologies you chose, assuming they worked as predicted. With the three I listed I can think about robots that never need to recharge other than “sunbathing”. 
  3. Apply the “one miracle rule”. In my example it would take a “miracle” to have powerful enough batteries that fit a humanoid robot. It would take a second “miracle” to get solar panels as efficient as necessary. Given this idea requires two miracles, I would not bet on it materializing. 
  4. Iterate. If I re-work the robot idea to remove at least one miracle, maybe defining a pre-fabricated house (static and large) as a “robot”, that would drop the number of miracles to one (in this case, people who can afford it, buying buying a pre-fab home) and that might materialize. 

Nothing that comes out of this exercise is easy to build, or has any guarantees of success. But by picking the hard problems I found my people. Colleagues, advisors, mentors, and leaders I follow, and people I help. 

………………………………………………………………………………………………………….

Ivan Santa Maria Filho has a BSc and MSc in computer science and a wide variety of experiences as individual contributor and manager, having owned a small software company and worked on multiple billion dollar products and services at Microsoft, Meta and Google. His main areas of expertise include vertical integration of stateful, large scale services with ephemeral VM infrastructure, and the infrastructure itself. Ivan Santa Maria Filho has a BSc and MSc in computer science and a wide variety of experiences as individual contributor and manager, having owned a small software company and worked on multiple billion dollar products and services at Microsoft, Meta and Google. His main areas of expertise include vertical integration of stateful, large scale services with ephemeral VM infrastructure, and the infrastructure itself.

(*)  Technical Architecture Focus: Scaling Pandas to Petabytes: The Architecture and Tradeoffs of BigQuery DataFrames. Interview with Ivan Santa Maria Filho, ODBMS Industry Watch, March 7, 2026

……………………………..

Follow us on X

Follow us on LinkedIn